Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / Physics Book Downloads / Math Methods in Physics Books / PDF Originals

Chow T.L. Mathematical methods for physicists.. a concise introduction (CUP, 2000)(569s)

PDF · 569 pages · 3.3 MB
Open PDF file

Published textbook by Tai L. Chow of California State University, Stanislaus, written for a two-semester intermediate course in mathematical physics. The table of contents covers vector and tensor analysis, ordinary differential equations, matrix algebra, Fourier series and integrals, linear vector spaces, complex variables, special functions, calculus of variations, Laplace transforms, partial differential equations and integral equations. It is a book by someone else, kept in the archive's collection of downloaded physics books.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Mathematical Methods for Physicists: A concise introduction CAMBRIDGE UNIVERSITY PRESSTAI L. CHOW Mathematical Methods for Ph/C121sicists A concise introduction This text is designed for an intermediate-level, two-semester undergraduate course in mathematical physics. It provides an accessible account of most of the current, important mathematical tools required in physics these days. It is assumed that the reader has an adequate preparation in general physics and calculus. The book bridges the gap between an introductory physics course and more advanced courses in classical mechanics, electricity and magnetism, quantum mechanics, and thermal and statistical physics. The text contains a large numberof worked examples to illustrate the mathematical techniques developed and to show their relevance to physics. The book is designed primarily for undergraduate physics majors, but could also be used by students in other subjects, such as engineering, astronomy and mathematics. TAI L . CHOW was born and raised in China. He received a BS degree in physics from the National Taiwan University, a Masters degree in physics from CaseWestern Reserve University, and a PhD in physics from the University of Rochester. Since 1969, Dr Chow has been in the Department of Physics at California State University, Stanislaus, and served as department chairman for17 years, until 1992. He served as Visiting Professor of Physics at University of California (at Davis and Berkeley) during his sabbatical years. He also worked as Summer Faculty Research Fellow at Stanford University and at NASA. Dr Chow has published more than 35 articles in physics journals and is the author of two textbooks and a solutions manual. PUBLISHED BY CAMBRIDGE UNIVERSITY PRESS (VIRTUAL PUBLISHING) FOR AND ON BEHALF OF THE PRESS SYNDICATE OF THE UNIVERSITY OF CAMBRIDGE The Pitt Building, Trumpington Street, Cambridge CB2 IRP 40 West 20th Street, New York, NY 10011-4211, USA 477 Williamstown Road, Port Melbourne, VIC 3207, Australia http://www.cambridge.org © Cambridge University Press 2000 This edition © Cambridge University Press (Virtual Publishing) 2003 First published in printed format 2000 A catalogue record for the original printed book is available from the British Library and from the Library of Congress Original ISBN 0 521 65227 8 hardback Original ISBN 0 521 65544 7 paperback ISBN 0 511 01022 2 virtual (netLibrary Edition) Mathematical Methods for Physicists A concise introduction TAI L . CHOW /C67alifornia State University /C67ontents Preface xv /C49 /C86ector and tensor anal/C121sis /C49 Vectors and scalars 1 Direction angles and direction cosines 3 Vector algebra 4 Equality of vectors 4Vector addition 4 Multiplication by a scalar 4 The scalar product 5 The vector (cross or outer) product 7 The triple scalar product /C65…/C66/C67†10 The triple vector product 11Change of coordinate system 11 The linear vector space /C86 n13 Vector di/C128erentiation 15Space curves 16 Motion in a plane 17 A vector treatment of classical orbit theory 18 Vector di/C128erential of a scalar field and the gradient 20 Conservative vector field 21The vector di/C128erential operator /C114 22 Vector di/C128erentiation of a vector field 22 The divergence of a vector 22 The operator /C114 2, the Laplacian 24 The curl of a vector 24 Formulas involving /C114 27 Orthogonal curvilinear coordinates 27 v Special orthogonal coordinate systems 32 Cylindrical coordinates …/C26; /C30; z†32 Spherical coordinates ( r; ;/C30 †34 Vector integration and integral theorems 35 Gauss’ theorem (the divergence theorem) 37 Continuity equation 39Stokes’ theorem 40 Green’s theorem 43 Green’s theorem in the plane 44 Helmholtz’s theorem 44Some useful integral relations 45Tensor analysis 47Contravariant and covariant vectors 48Tensors of second rank 48Basic operations with tensors 49Quotient law 50The line element and metric tensor 51Associated tensors 53Geodesics in a Riemannian space 53Covariant di/C128erentiation 55 Problems 57 /C50 /C79rdinar/C121 di/C128erential equations /C54/C50 First-order di/C128erential equations 63 Separable variables 63 Exact equations 67Integrating factors 69Bernoulli’s equation 72 Second-order equations with constant coecients 72 Nature of the solution of linear equations 73General solutions of the second-order equations 74Finding the complementary function 74Finding the particular integral 77Particular integral and the operator D…ˆd=dx†78 Rules for /C68operators 79 The Euler linear equation 83Solutions in power series 85 Ordinary and singular points of a di/C128erential equation 86Frobenius and Fuchs theorem 86 Simultaneous equations 93The gamma and beta functions 94 Problems 96CONTENTS vi /C51 Matri/C120 algebra /C49/C48/C48 Definition of a matrix 100 Four basic algebra operations for matrices 102 Equality of matrices 102Addition of matrices 102 Multiplication of a matrix by a number 103 Matrix multiplication 103 The commutator 107 Powers of a matrix 107 Functions of matrices 107 Transpose of a matrix 108Symmetric and skew-symmetric matrices 109 The matrix representation of a vector product 110 The inverse of a matrix 111 A method for finding ~A ÿ1112 Systems of linear equations and the inverse of a matrix 113Complex conjugate of a matrix 114 Hermitian conjugation 114 Hermitian/anti-hermitian matrix 114 Orthogonal matrix (real) 115 Unitary matrix 116 Rotation matrices 117Trace of a matrix 121 Orthogonal and unitary transformations 121 Similarity transformation 122 The matrix eigenvalue problem 124 Determination of eigenvalues and eigenvectors 124 Eigenvalues and eigenvectors of hermitian matrices 128 Diagonalization of a matrix 129 Eigenvectors of commuting matrices 133 Cayley–Hamilton theorem 134Moment of inertia matrix 135 Normal modes of vibrations 136 Direct product of matrices 139 Problems 140 /C52 Fourier series and integrals /C49/C52/C52 Periodic functions 144 Fourier series; Euler–Fourier formulas 146 Gibb’s phenomena 150 Convergence of Fourier series and Dirichlet conditions 150CONTENTS vii Half-range Fourier series 151 Change of interval 152Parseval’s identity 153 Alternative forms of Fourier series 155 Integration and di/C128erentiation of a Fourier series 157 Vibrating strings 157 The equation of motion of transverse vibration 157Solution of the wave equation 158 /C82L/C67 circuit 160 Orthogonal functions 162 Multiple Fourier series 163Fourier integrals and Fourier transforms 164 Fourier sine and cosine transforms 172 Heisenberg’s uncertainty principle 173 Wave packets and group velocity 174 Heat conduction 179 Heat conduction equation 179 Fourier transforms for functions of several variables 182 The Fourier integral and the delta function 183 Parseval’s identity for Fourier integrals 186 The convolution theorem for Fourier transforms 188 Calculations of Fourier transforms 190The delta function and Green’s function method 192 Problems 195 /C53 Linear /C118ector spaces /C49/C57/C57 Euclidean n-space /C69 n199 General linear vector spaces 201 Subspaces 203 Linear combination 204Linear independence, bases, and dimensionality 204 Inner product spaces (unitary spaces) 206 The Gram–Schmidt orthogonalization process 209 The Cauchy–Schwarz inequality 210 Dual vectors and dual spaces 211 Linear operators 212 Matrix representation of operators 214 The algebra of linear operators 215 Eigenvalues and eigenvectors of an operator 217 Some special operators 217 The inverse of an operator 218CONTENTS viii The adjoint operators 219 Hermitian operators 220Unitary operators 221 The projection operators 222 Change of basis 224Commuting operators 225 Function spaces 226 Problems 230 /C54 Functions of a comple/C120 /C118ariable /C50/C51/C51 Complex numbers 233 Basic operations with complex numbers 234 Polar form of complex number 234 De Moivre’s theorem and roots of complex numbers 237 Functions of a complex variable 238Mapping 239 Branch lines and Riemann surfaces 240 The di/C128erential calculus of functions of a complex variable 241 Limits and continuity 241Derivatives and analytic functions 243The Cauchy–Riemann conditions 244 Harmonic functions 247 Singular points 248 Elementary functions of z249 The exponential functions e z(or exp( z)†249 Trigonometric and hyperbolic functions 251The logarithmic functions /C119ˆlnz252 Hyperbolic functions 253 Complex integration 254 Line integrals in the complex plane 254 Cauchy’s integral theorem 257 Cauchy’s integral formulas 260 Cauchy’s integral formulas for higher derivatives 262 Series representations of analytic functions 265 Complex sequences 265 Complex series 266 Ratio test 268 Uniform covergence and the Weierstrass M/C45test 268 Power series and Taylor series 269Taylor series of elementary functions 272Laurent series 274CONTENTS ix Integration by the method of residues 279 Residues 279 The residue theorem 282 Evaluation of real definite integrals 283 Improper integrals of the rational functionZ1 ÿ1f…x†dx 283 Integrals of the rational functions of sin and cos Z2 0/C71…sin;cos†d286 Fourier integrals of the formZ1 ÿ1f…x†sinmx cosmx/C26/C27 dx 288 Problems 292 /C55 /C83pecial functions of mathematical ph/C121sics /C50/C57/C54 Legendre’s equation 296 Rodrigues’ formula for Pn…x†299 The generating function for Pn…x†301 Orthogonality of Legendre polynomials 304 The associated Legendre functions 307 Orthogonality of associated Legendre functions 309 Hermite’s equation 311 Rodrigues’ formula for Hermite polynomials Hn…x†313 Recurrence relations for Hermite polynomials 313Generating function for the H n…x†314 The orthogonal Hermite functions 314 Laguerre’s equation 316 The generating function for the Laguerre polynomials Ln…x†317 Rodrigues’ formula for the Laguerre polynomials Ln…x†318 The orthogonal Laugerre functions 319 The associated Laguerre polynomials Lm n…x†320 Generating function for the associated Laguerre polynomials 320 Associated Laguerre function of integral order 321 Bessel’s equation 321 Bessel functions of the second kind Yn…x†325 Hanging flexible chain 328Generating function for /C74 n…x†330 Bessel’s integral representation 331Recurrence formulas for /C74 n…x†332 Approximations to the Bessel functions 335Orthogonality of Bessel functions 336 Spherical Bessel functions 338CONTENTS x Sturm–Liouville systems 340 Problems 343 /C56 /C84he calculus of /C118ariations /C51/C52/C55 The Euler–Lagrange equation 348 Variational problems with constraints 353Hamilton’s principle and Lagrange’s equation of motion 355 Rayleigh–Ritz method 359 Hamilton’s principle and canonical equations of motion 361 The modified Hamilton’s principle and the Hamilton–Jacobi equation 364 Variational problems with several independent variables 367 Problems 369 /C57 /C84he Laplace transformation /C51/C55/C50 Definition of the Lapace transform 372 Existence of Laplace transforms 373 Laplace transforms of some elementary functions 375 Shifting (or translation) theorems 378 The first shifting theorem 378The second shifting theorem 379 The unit step function 380 Laplace transform of a periodic function 381 Laplace transforms of derivatives 382 Laplace transforms of functions defined by integrals 383 A note on integral transformations 384 Problems 385 /C49/C48 Partial di/C128erential equations /C51/C56/C55 Linear second-order partial di/C128erential equations 388 Solutions of Laplace’s equation: separation of variables 392 Solutions of the wave equation: separation of variables 402 Solution of Poisson’s equation. Green’s functions 404 Laplace transform solutions of boundary-value problems 409 Problems 410 /C49/C49 /C83imple linear integral equations /C52/C49/C51 Classification of linear integral equations 413 Some methods of solution 414 Separable kernel 414Neumann series solutions 416CONTENTS xi Transformation of an integral equation into a di/C128erential equation 419 Laplace transform solution 420Fourier transform solution 421 The Schmidt–Hilbert method of solution 421Relation between di/C128erential and integral equations 425 Use of integral equations 426 Abel’s integral equation 426Classical simple harmonic oscillator 427 Quantum simple harmonic oscillator 427 Problems 428 /C49/C50 /C69lements of group theor/C121 /C52/C51/C48 Definition of a group (group axioms) 430 Cyclic groups 433 Group multiplication table 434 Isomorphic groups 435 Group of permutations and Cayley’s theorem 438Subgroups and cosets 439 Conjugate classes and invariant subgroups 440 Group representations 442 Some special groups 444 The symmetry group D 2;D3446 One-dimensional unitary group U…1†449 Orthogonal groups SO…2†andSO…3†450 TheSU…n†groups 452 Homogeneous Lorentz group 454 Problems 457 /C49/C51 Numerical methods /C52/C53/C57 Interpolation 459 Finding roots of equations 460 Graphical methods 460 Method of linear interpolation (method of false position) 461 Newton’s method 464 Numerical integration 466 The rectangular rule 466The trapezoidal rule 467 Simpson’s rule 469 Numerical solutions of di/C128erential equations 469 Euler’s method 470 The three-term Taylor series method 472CONTENTS xii The Runge–Kutta method 473 Equations of higher order. System of equations 476 Least-squares fit 477 Problems 478 /C49/C52 Introduction to probabilit/C121 theor/C121 /C52/C56/C49 A definition of probability 481Sample space 482 Methods of counting 484 Permutations 484 Combinations 485 Fundamental probability theorems 486 Random variables and probability distributions 489 Random variables 489Probability distributions 489 Expectation and variance 490 Special probability distributions 491 The binomial distribution 491 The Poisson distribution 495 The Gaussian (or normal) distribution 497 Continuous distributions 500 The Gaussian (or normal) distribution 502The Maxwell–Boltzmann distribution 503 Problems 503 /C65ppendi/C120 /C49 Preliminaries (review of fundamental concepts) /C53/C48/C54 Inequalities 507 Functions 508 Limits 510 Infinite series 511 Tests for convergence 513Alternating series test 516 Absolute and conditional convergence 517 Series of functions and uniform convergence 520 Weistrass Mtest 521 Abel’s test 522Theorem on power series 524 Taylor’s expansion 524 Higher derivatives and Leibnitz’s formula for nth derivative of a product 528 Some important properties of definite integrals 529CONTENTS xiii Some useful methods of integration 531 Reduction formula 533Di/C128erentiation of integrals 534 Homogeneous functions 535 Taylor series for functions of two independent variables 535 Lagrange multiplier 536 /C65ppendi/C120 /C50 /C68eterminants /C53/C51/C56 Determinants, minors, and cofactors 540Expansion of determinants 541 Properties of determinants 542Derivative of a determinant 547 /C65ppendi/C120 /C51 /C84able of function /C70…x†ˆ 1  2pZx 0eÿt2=2dt/C53/C52/C56 /C70urther reading 549 Index 551CONTENTS xiv Preface This book evolved from a set of lecture notes for a course on ‘Introduction to Mathematical Physics’, that I have given at California State University, Stanislaus (CSUS) for many years. Physics majors at CSUS take introductory mathematical physics before the physics core courses, so that they may acquire the expected level of mathematical competency for the core course. It is assumed that the student has an adequate preparation in general physics and a good understanding of the mathematical manipulations of calculus. For the student who is in need of a review of calculus, however, Appendix 1 and Appendix 2 are included. This book is not encyclopedic in character, nor does it give in a highly mathe- matical rigorous account. Our emphasis in the text is to provide an accessibleworking knowledge of some of the current important mathematical tools required in physics. The student will find that a generous amount of detail has been given mathe- matical manipulations, and that ‘it-may-be-shown-thats’ have been kept to a minimum. However, to ensure that the student does not lose sight of the develop- ment underway, some of the more lengthy and tedious algebraic manipulations have been omitted when possible. Each chapter contains a number of physics examples to illustrate the mathe- matical techniques just developed and to show their relevance to physics. They supplement or amplify the material in the text, and are arranged in the order in which the material is covered in the chapter. No e/C128ort has been made to trace theorigins of the homework problems and examples in the book. A solution manual for instructors is available from the publishers upon adoption. Many individuals have been very helpful in the preparation of this text. I wish to thank my colleagues in the physics department at CSUS. Any suggestions for improvement of this text will be greatly appreciated. Turlock/C44 /C67alifornia TAI L . CHOW 2000 xv 1 /C86ector and tensor analysis /C86ectors and scalars Vector methods have become standard tools for the physicists. In this chapter we discuss the properties of the vectors and vector fields that occur in classical physics. We will do so in a way, and in a notation, that leads to the formation of abstract linear vector spaces in Chapter 5. A physical quantity that is completely specified, in appropriate units, by a single number (called its magnitude) such as volume, mass, and temperature is called a scalar. Scalar quantities are treated as ordinary real numbers. They obey all the regular rules of algebraic addition, subtraction, multiplication, division, and soon. There are also physical quantities which require a magnitude and a direction for their complete specification. These are called vectors iftheir combination with each other is commutative (that is the order of addition may be changed withouta/C128ecting the result). Thus not all quantities possessing magnitude and direction are vectors. Angular displacement, for example, may be characterised by magni- tude and direction but is not a vector, for the addition of two or more angular displacements is not, in general, commutative (Fig. 1.1). In print, we shall denote vectors by boldface letters (such as /C65) and use ordin- ary italic letters (such as A) for their magnitudes; in writing, vectors are usually represented by a letter with an arrow above it such as /C126A. A given vector /C65(or/C126A) can be written as /C65ˆA^A; …1:1† where Ais the magnitude of vector /C65and so it has unit and dimension, and ^Ais a dimensionless unit vector with a unity magnitude having the direction of /C65. Thus ^Aˆ/C65=A. 1 A vector quantity may be represented graphically by an arrow-tipped line seg- ment. The length of the arrow represents the magnitude of the vector, and the direction of the arrow is that of the vector, as shown in Fig. 1.2. Alternatively, avector can be specified by its components (projections along the coordinate axes) and the unit vectors along the coordinate axes (Fig. 1.3): /C65ˆA 1^e1‡A2^e2‡A^e3ˆX3 iˆ1Ai^ei; …1:2† where ^ei(iˆ1;2;3) are unit vectors along the rectangular axes xi…x1ˆx;x2ˆy; x3ˆz†; they are normally written as ^i;^j;^kin general physics textbooks. The component triplet ( A1;A2;A3) is also often used as an alternate designation for vector /C65: /C65ˆ…A1;A2;A3†: …1:2a† This algebraic notation of a vector can be extended (or generalized) to spaces of dimension greater than three, where an ordered n-tuple of real numbers, (A1;A2;...;An), represents a vector. Even though we cannot construct physical vectors for n/C623, we can retain the geometrical language for these n-dimensional generalizations. Such abstract ‘‘vectors’’ will be the subject of Chapter 5. 2VECTOR AND TENSOR ANALYSIS Figure 1.1. Rotation of a parallelpiped about coordinate axes. Figure 1.2. Graphical representation of vector /C65. /C68irection angles and direction cosines We can express the unit vector ^Ain terms of the unit coordinate vectors ^ei.F r o m Eq. (1.2), /C65ˆA1^e1‡A2^e2‡A^e3, we have /C65ˆAA1 A^e1‡A2 A^e2‡A3 A^e3 ˆA^A: Now A1=Aˆcos ;A2=Aˆcos/C12,a n d A3=Aˆcos/C13are the direction cosines of the vector /C65, and ,/C12,a n d /C13are the direction angles (Fig. 1.4). Thus we can write /C65ˆA…cos ^e1‡cos/C12^e2‡cos/C13^e3†ˆA^A; it follows that ^Aˆ…cos ^e1‡cos/C12^e2‡cos/C13^e3†ˆ… cos ;cos/C12;cos/C13†: …1:3† 3DIRECTION ANGLES AND DIRECTION COSINES Figure 1.3. A vector /C65in Cartesian coordinates. Figure 1.4. Direction angles of vector /C65. /C86ector algebra Equality of vectors Two vectors, say /C65and/C66, are equal if, and only if, their respective components are equal: /C65ˆ/C66or…A1;A2;A3†ˆ…B1;B2;B3† is equivalent to the three equations A1ˆB1;A2ˆB2;A3ˆB3: Geometrically, equal vectors are parallel and have the same length, but do not necessarily have the same position. /C86ector addition The addition of two vectors is defined by the equation /C65‡/C66ˆ…A1;A2;A3†‡…B1;B2;B3†ˆ…A1‡B1;A2‡B2;A3‡B3†: That is, the sum of two vectors is a vector whose components are sums of thecomponents of the two given vectors. We can add two non-parallel vectors by graphical method as shown in Fig. 1.5. To add vector /C66to vector /C65, shift /C66parallel to itself until its tail is at the head of /C65. The vector sum /C65‡/C66is a vector /C67drawn from the tail of /C65to the head of /C66. The order in which the vectors are added does not a/C128ect the result. /C77ultiplication by a scalar Ifcis scalar then c/C65ˆ…cA 1;cA2;cA3†: Geometrically, the vector c/C65is parallel to /C65and is ctimes the length of /C65. When cˆÿ1, the vector ÿ/C65is one whose direction is the reverse of that of /C65, but both 4VECTOR AND TENSOR ANALYSIS Figure 1.5. Addition of two vectors. have the same length. Thus, subtraction of vector /C66from vector /C65is equivalent to adding ÿ/C66to/C65: /C65ÿ/C66ˆ/C65‡… ÿ /C66†: We see that vector addition has the following properties: (a)/C65‡/C66ˆ/C66‡/C65 (commutativity); (b) (/C65‡/C66†‡/C67ˆ/C65‡…/C66‡/C67† (associativity); (c)/C65‡/C48ˆ/C48‡/C65ˆ/C65; (d)/C65‡… ÿ /C65†ˆ/C48: We now turn to vector multiplication. Note that division by a vector is not defined: expressions such as k=/C65or/C66=/C65are meaningless. There are several ways of multiplying two vectors, each of which has a special meaning; two types are defined. /C84he scalar product The scalar (dot or inner) product of two vectors /C65and/C66is a real number defined (in geometrical language) as the product of their magnitude and the cosine of the (smaller) angle between them (Figure 1.6): /C65/C66ABcos …0†: …1:4† It is clear from the definition (1.4) that the scalar product is commutative: /C65/C66ˆ/C66/C65; …1:5† and the product of a vector with itself gives the square of the dot product of thevector: /C65/C65ˆA 2: …1:6† If/C65/C66ˆ0 and neither /C65nor/C66is a null (zero) vector, then /C65is perpendicular to /C66. 5THE SCALAR PRODUCT Figure 1.6. The scalar product of two vectors. We can get a simple geometric interpretation of the dot product from an inspection of Fig. 1.6: …Bcos†Aˆprojection of /C66onto /C65multiplied by the magnitude of /C65; …Acos†Bˆprojection of /C65onto /C66multiplied by the magnitude of /C66: If only the components of /C65and/C66are known, then it would not be practical to calculate /C65/C66from definition (1.4). But, in this case, we can calculate /C65/C66in terms of the components: /C65/C66ˆ…A1^e1‡A2^e2‡A3^e3†…B1^e1‡B2^e2‡B3^e3†; …1:7† the right hand side has nine terms, all involving the product ^ei^ej. Fortunately, the angle between each pair of unit vectors is 90 8, and from (1.4) and (1.6) we find that ^ei^ejˆij; i;jˆ1;2;3; …1:8† where ijis the Kronecker delta symbol ijˆ0;ifi6ˆj; 1;ifiˆj:( …1:9† After we use (1.8) to simplify the resulting nine terms on the right-side of (7), we obtain /C65/C66ˆA1B1‡A2B2‡A3B3ˆX3 iˆ1AiBi: …1:10† The law of cosines for plane triangles can be easily proved with the application of the scalar product: refer to Fig. 1.7, where /C67is the resultant vector of /C65and/C66. Taking the dot product of /C67with itself, we obtain C2ˆ/C67/C67ˆ…/C65‡/C66†…/C65‡/C66† ˆA2‡B2‡2/C65/C66ˆA2‡B2‡2ABcos; which is the law of cosines. 6VECTOR AND TENSOR ANALYSIS Figure 1.7. Law of cosines. A simple application of the scalar product in physics is the work /C87done by a constant force F:/C87ˆFr, where ris the displacement vector of the object moved by F. /C84he /C118ector (cross or outer) product The vector product of two vectors /C65and/C66is a vector and is written as /C67ˆ/C65/C66: …1:11† As shown in Fig. 1.8, the two vectors /C65and/C66form two sides of a parallelogram. We define /C67to be perpendicular to the plane of this parallelogram with its magnitude equal to the area of the parallelogram. And we choose the direction of/C67along the thumb of the right hand when the fingers rotate from /C65to/C66(angle of rotation less than 180 8). /C67ˆ/C65/C66ˆABsin^eC …0†: …1:12† From the definition of the vector product and following the right hand rule, we can see immediately that /C65/C66ˆÿ/C66/C65: …1:13† Hence the vector product is not commutative. If /C65and/C66are parallel, then it follows from Eq. (1.12) that /C65/C66ˆ0: …1:14† In particular /C65/C65ˆ0: …1:14a† In vector components, we have /C65/C66ˆ…A1^e1‡A2^e2‡A3^e3†…B1^e1‡B2^e2‡B3^e3†: …1:15† 7THE VECTOR (CROSS OR OUTER) PRODUCT Figure 1.8. The right hand rule for vector product. Using the following relations ^ei^eiˆ0;iˆ1;2;3; ^e1^e2ˆ^e3;^e2^e3ˆ^e1;^e3^e1ˆ^e2;…1:16† Eq. (1.15) becomes /C65/C66ˆ…A2B3ÿA3B2†^e1‡…A3B1ÿA1B3†^e2‡…A1B2ÿA2B1†^e3:…1:15a† This can be written as an easily remembered determinant of third order: /C65/C66ˆ^e1^e2^e3 A1A2A3 B1B2B3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: …1:17† The expansion of a determinant of third order can be obtained by diagonal multi- plication by repeating on the right the first two columns of the determinant and adding the signed products of the elements on the various diagonals in the result-ing array: The non-commutativity of the vector product of two vectors now appears as a consequence of the fact that interchanging two rows of a determinant changes itssign, and the vanishing of the vector product of two vectors in the same direction appears as a consequence of the fact that a determinant vanishes if one of its rows is a multiple of another. The determinant is a basic tool used in physics and engineering. The reader is assumed to be familiar with this subject. Those who are in need of review should read Appendix II. The vector resulting from the vector product of two vectors is called an axial vector, while ordinary vectors are sometimes called polar vectors. Thus, in Eq.(1.11), /C67is a pseudovector, while /C65and/C66are axial vectors. On an inversion of coordinates, polar vectors change sign but an axial vector does not change sign. A simple application of the vector product in physics is the torque /C115of a force F about a point O:/C115ˆFr, where ris the vector from Oto the initial point of the force F(Fig. 1.9). We can write the nine equations implied by Eq. (1.16) in terms of permutation symbols /C34 ijk: ^ei^ejˆ/C34ijk^ek; …1:16a† 8VECTOR AND TENSOR ANALYSIS a1a2a3 b1b2b3 c1c2cc2 435a 1a2 b1b2 c1c2 ÿÿÿ ‡‡‡ ------ ------ ------ÿ ÿ !ÿ ÿ !ÿ ÿ ! where /C34ijkis defined by /C34ijkˆ‡1 ÿ1 0if…i;j;k†is an even permutation of …1;2;3†; if…i;j;k†is an odd permutation of …1;2;3†; otherwise …for example ;if 2 or more indices are equal †:8 < :…1:18† It follows immediately that /C34ijkˆ/C34kijˆ/C34jkiˆÿ/C34jikˆÿ/C34kjiˆÿ/C34ikj: There is a very useful identity relating the /C34ijkand the Kronecker delta symbol: X3 kˆ1/C34mnk/C34ijkˆminjÿmjni; …1:19† X j;k/C34mjk/C34njkˆ2mn;X i;j;k/C342 ijkˆ6: …1:19a† Using permutation symbols, we can now write the vector product /C65/C66as /C65/C66ˆX3 iˆ1Ai^ei/C32! X3 jˆ1Bj^ej/C32! ˆX3 i;jAiBj^ei^ejÿ ˆX3 i;j;kAiBj/C34ijkÿ^ek: Thus the kth component of /C65/C66is …/C65/C66†kˆX i;jAiBj/C34ijkˆX i;j/C34kijAiBj: Ifkˆ1, we obtain the usual geometrical result: …/C65/C66†1ˆX i;j/C341ijAiBjˆ/C34123A2B3‡/C34132A3B2ˆA2B3ÿA3B2: 9THE VECTOR (CROSS OR OUTER) PRODUCT Figure 1.9. The torque of a force about a point O. /C84he triple scalar product /C65 E(/C66/C67) We now briefly discuss the scalar /C65…/C66/C67†. This scalar represents the volume of the parallelepiped formed by the coterminous sides /C65,/C66,/C67, since /C65…/C66/C67†ˆABC sincos ˆ/C104Sˆvolume ; Sbeing the area of the parallelogram with sides /C66and/C67, and hthe height of the parallelogram (Fig. 1.10). Now /C65…/C66/C67†ˆA1^e1‡A2^e2‡A3^e3 …† ^e1^e2^e3 B1B2B3 C1C2C3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 ˆA 1…B2C3ÿB3C2†‡A2…B3C1ÿB1C3†‡A3…B1C2ÿB2C1† so that /C65…/C66/C67†ˆA1A2A3 B1B2B3 C1C2C3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: …1:20† The exchange of two rows (or two columns) changes the sign of the determinant but does not change its absolute value. Using this property, we find /C65…/C66/C67†ˆA 1A2A3 B1B2B3 C1C2C3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆÿC 1C2C3 B1B2B3 A1A2A3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ/C67…/C65/C66†; that is, the dot and the cross may be interchanged in the triple scalar product. /C65…/C66/C67†ˆ… /C65/C66†/C67 …1:21† 10VECTOR AND TENSOR ANALYSIS Figure 1.10. The triple scalar product of three vectors /C65,/C66,/C67. In fact, as long as the three vectors appear in cyclic order, /C65!/C66!/C67!/C65, then the dot and cross may be inserted between any pairs: /C65…/C66/C67†ˆ/C66…/C67/C65†ˆ/C67…/C65/C66†: It should be noted that the scalar resulting from the triple scalar product changes sign on an inversion of coordinates. For this reason, the triple scalar product is sometimes called a pseudoscalar. /C84he triple /C118ector product The triple product /C65…/C66/C67) is a vector, since it is the vector product of two vectors: /C65and/C66/C67. This vector is perpendicular to /C66/C67and so it lies in the plane of /C66and/C67.I f/C66is not parallel to /C67,/C65…/C66/C67†ˆx/C66‡y/C67. Now dot both sides with /C65and we obtain x…/C65/C66†‡y…/C65/C67†ˆ0, since /C65‰/C65…/C66/C67†Š ˆ0. Thus x=…/C65/C67†ˆÿ y=…/C65/C66†…is a scalar † and so /C65…/C66/C67†ˆx/C66‡y/C67ˆ‰/C66…/C65/C67†ÿ/C67…/C65/C66†Š: We now show that ˆ1. To do this, let us consider the special case when /C66ˆ/C65. Dot the last equation with /C67: /C67‰/C65…/C65/C67†Š ˆ‰…/C65/C67†2ÿ/C652/C672Š; or, by an interchange of dot and cross ÿ…/C65/C67†2ˆ‰…/C65/C67†2ÿ/C652/C672Š: In terms of the angles between the vectors and their magnitudes the last equationbecomes ÿA 2C2sin2ˆ…A2C2cos2ÿA2C2†ˆÿ A2C2sin2; hence ˆ1. And so /C65…/C66/C67†ˆ/C66…/C65/C67†ÿ/C67…/C65/C66†: …1:22† /C67hange of coordinate s/C121stem Vector equations are independent of the coordinate system we happen to use. But the components of a vector quantity are di/C128erent in di/C128erent coordinate systems. We now make a brief study of how to represent a vector in di/C128erent coordinate systems. As the rectangular Cartesian coordinate system is the basic type of coordinate system, we shall limit our discussion to it. Other coordinate systems 11THE TRIPLE VECTOR PRODUCT will be introduced later. Consider the vector /C65expressed in terms of the unit coordinate vectors …^e1;^e2;^e3†: /C65ˆA1^e1‡A2^e2‡A^e3ˆX3 iˆ1Ai^ei: Relative to a new system …^e0 1;^e0 2;^e0 3†that has a di/C128erent orientation from that of the old system …^e1;^e2;^e3†, vector /C65is expressed as /C65ˆA0 1^e0 1‡A0 2^e0 2‡A0^e0 3ˆX3 iˆ1A0 i^e0 i: Note that the dot product /C65^e0 1is equal to A0 1, the projection of /C65on the direction of^e0 1;/C65^e0 2is equal to A0 2, and /C65^e0 3is equal to A0 3. Thus we may write A0 1ˆ…^e1^e0 1†A1‡…^e2^e0 1†A2‡…^e3^e0 1†A3; A0 2ˆ…^e1^e0 2†A1‡…^e2^e0 2†A2‡…^e3^e0 2†A3; A0 3ˆ…^e1^e0 3†A1‡…^e2^e0 3†A2‡…^e3^e0 3†A3:9 >>= >>;…1:23† The dot products …^ei^e0 j†are the direction cosines of the axes of the new coordi- nate system relative to the old system: ^e0 i^ejˆcos…x0 i;xj†; they are often called the coecients of transformation. In matrix notation, we can write the above system of equations as A0 1 A0 2 A0 30 B@1 CAˆ^e1^e0 1^e2^e0 1^e3^e0 1 ^e1^e0 2^e2^e0 2^e3^e0 2 ^e1^e0 3^e2^e0 3^e3^e0 30 B@1 CAA1 A2 A30 B@1 CA: The 3 3 matrix in the above equation is called the rotation (or transformation) matrix, and is an orthogonal matrix. One advantage of using a matrix is that successive transformations can be handled easily by means of matrix multiplica- tion. Let us digress for a quick review of some basic matrix algebra. A full account of matrix method is given in Chapter 3. A matrix is an ordered array of scalars that obeys prescribed rules of addition and multiplication. A particular matrix element is specified by its row numberfollowed by its column number. Thus a ijis the matrix element in the ith row and jth column. Alternative ways of representing matrix ~Aare /C91aij/C93 or the entire array ~Aˆa11a12:::a1n a21a22:::a2n ::: ::: ::: ::: am1am2:::amn0 BBBB@1 CCCCA: 12VECTOR AND TENSOR ANALYSIS ~Ais an nmmatrix. A vector is represented in matrix form by writing its components as either a row or column array, such as ~Bˆ…b11b12b13†or ~Cˆc11 c21 c310 B@1 CA; where b11ˆbx;b12ˆby;b13ˆbz, and c11ˆcx;c21ˆcy;c31ˆcz. The multiplication of a matrix ~Aand a matrix ~Bis defined only when the number of columns of ~Ais equal to the number of rows of ~B, and is performed in the same way as the multiplication of two determinants: if ~C/C61~A~B, then cijˆX kaikbk/C108: We illustrate the multiplication rule for the case of the 3 3 matrix ~Amultiplied by the 3 3 matrix ~B: If we denote the direction cosines ^e0 i^ejbyij, then Eq. (1.23) can be written as A0 iˆX3 jˆ1^e0 i^ejAjˆX3 jˆ1ijAj: …1:23a† It can be shown (Problem 1.9) that the quantities ijsatisfy the following relations X3 iˆ1ijikˆjk…j;kˆ1;2;3†: …1:24† Any linear transformation, such as Eq. (1.23a), that has the properties required by Eq. (1.24) is called an orthogonal transformation, and Eq. (1.24) is known as the orthogonal condition. /C84he linear /C118ector space /C86n We have found that it is very convenient to use vector components, in particular,the unit coordinate vectors ^e i(iˆ1, 2, 3). The three unit vectors ^eiare orthogonal and normal, or, as we shall say, orthonormal. This orthonormal propertyis conveniently written as Eq. (1.8). But there is nothing special about these 13THE LINEAR VECTOR SPACE /C86n . orthonormal unit vectors ^ei. If we refer the components of the vectors to a di/C128erent system of rectangular coordinates, we need to introduce another set of three orthonormal unit vectors ^f1;^f2, and ^f3: ^fi^fjˆij…i;jˆ1;2;3†: …1:8a† For any vector /C65we now write /C65ˆX3 iˆ1ci^fi;and ciˆ^fi/C65: We see that we can define a large number of di/C128erent coordinate systems. But the physically significant quantities are the vectors themselves and certain func-tions of these, which are independent of the coordinate system used. The ortho- normal condition (1.8) or (1.8a) is convenient in practice. If we also admit oblique Cartesian coordinates then the ^f ineed neither be normal nor orthogonal; they could be any three non-coplanar vectors, and any vector /C65can still be written as a linear superposition of the ^fi /C65ˆc1^f1‡c2^f2‡c3^f3: …1:25† Starting with the vectors ^fi, we can find linear combinations of them by the algebraic operations of vector addition and multiplication of vectors by scalars,and then the collection of all such vectors makes up the three-dimensional linear space often called /C86 3(V for vector) or R3(/C82for real) or /C693(Efor Euclidean). The vectors ^f1;^f2;^f3are called the base vectors or bases of the vector space /C863. Any set of vectors, such as the ^fi, which can serve as the bases or base vectors of /C863is called complete, and we say it spans the linear vector space. The base vectors arealso linearly independent because no relation of the form c 1^f1‡c2^f2‡c3^f3ˆ0 …1:26† exists between them, unless c1ˆc2ˆc3ˆ0. The notion of a vector space is much more general than the real vector space /C863. Extending the concept of /C863, it is convenient to call an ordered set of n matrices, or functions, or operators, a ‘vector’ (or an n-vector) in the n-dimen- sional space /C86n. Chapter 5 will provide justification for doing this. Taking a cue from /C863, vector addition in /C86nis defined to be …x1;...;xn†‡…y1;...;yn†ˆ…x1‡y1;...;xn‡yn†… 1:27† and multiplication by scalars is defined by …x1;...;xn†ˆ… x1;...; xn†; …1:28† 14VECTOR AND TENSOR ANALYSIS where is real. With these two algebraic operations of vector addition and multi- plication by scalars, we call /C86na vector space. In addition to this algebraic structure, /C86nhas geometric structure derived from the length defined to be Xn jˆ1x2 j/C32!1=2 ˆ x2 1‡‡ x2nq …1:29† The dot product of two n-vectors can be defined by …x1;...;xn†…y1;...;yn†ˆXn jˆ1xjyj: …1:30† In/C86n, vectors are not directed line segments as in /C863; they may be an ordered set ofnoperators, matrices, or functions. We do not want to become sidetracked from our main goal of this chapter, so we end our discussion of vector space here. /C86ector di/C128erentiation Up to this point we have been concerned mainly with vector algebra. A vector may be a function of one or more scalars and vectors. We have encountered, for example, many important vectors in mechanics that are functions of time and position variables. We now turn to the study of the calculus of vectors. Physicists like the concept of field and use it to represent a physical quantity that is a function of position in a given region. Temperature is a scalar field, because its value depends upon location: to each point ( x,y,z) is associated a temperature T…x;y;z†. The function T…x;y;z†is a scalar field, whose value is a real number depending only on the point in space but not on the particular choiceof the coordinate system. A vector field, on the other hand, associates with each point a vector (that is, we associate three numbers at each point), such as the wind velocity or the strength of the electric or magnetic field. When described in a rotated system, for example, the three components of the vector associated with one and the same point will change in numerical value. Physically and geo- metrically important concepts in connection with scalar and vector fields are the gradient, divergence, curl, and the corresponding integral theorems. The basic concepts of calculus, such as continuity and di/C128erentiability, can be naturally extended to vector calculus. Consider a vector /C65, whose components are functions of a single variable u. If the vector /C65represents position or velocity, for example, then the parameter uis usually time t, but it can be any quantity that determines the components of /C65. If we introduce a Cartesian coordinate system, the vector function /C65(u) may be written as /C65…u†ˆA 1…u†^e1‡A2…u†^e2‡A3…u†^e3: …1:31† 15VECTOR DIFFERENTIATION /C65(u) is said to be continuous at uˆu0if it is defined in some neighborhood of u0and lim u!u0A…u†ˆA…u0†: …1:32† Note that /C65(u) is continuous at u0if and only if its three components are con- tinuous at u0. /C65(u) is said to be di/C128erentiable at a point uif the limit d/C65…u† duˆlim u!0/C65…u‡u†ÿ/C65…u† u…1:33† exists. The vector /C650…u†ˆd/C65…u†=duis called the derivative of /C65(u); and to di/C128er- entiate a vector function we di/C128erentiate each component separately: /C650…u†ˆA0 1…u†^e1‡A0 2…u†^e2‡A0 3…u†^e3: …1:33a† Note that the unit coordinate vectors are fixed in space. Higher derivatives of /C65(u) can be similarly defined. If/C65is a vector depending on more than one scalar variable, say u,/C118for example, we write /C65ˆ/C65…u;/C118†. Then d/C65ˆ…/C64/C65=/C64u†du‡…/C64/C65=/C64/C118†d/C118 …1:34† is the di/C128erential of /C65, and /C64/C65 /C64uˆlim u!0/C65…u‡u;/C118†ÿ/C65…u;/C118† /C64u…1:34a† and similarly for /C64/C65=/C64/C118. Derivatives of products obey rules similar to those for scalar functions. However, when cross products are involved the order may be important. /C83pace cur/C118es As an application of vector di/C128erentiation, let us consider some basic facts about curves in space. If /C65(u) is the position vector r(u) joining the origin of a coordinate system and any point P…x1;x2;x3†in space as shown in Fig. 1.11, then Eq. (1.31) becomes r…u†ˆx1…u†^e1‡x2…u†^e2‡x3…u†^e3: …1:35† Asuchanges, the terminal point Pofrdescribes a curve /C67in space. Eq. (1.35) is called a parametric representation of the curve /C67, and uis the parameter of this representation. Then r uˆr…u‡u†ÿr…u† u 16VECTOR AND TENSOR ANALYSIS is a vector in the direction of r, and its limit (if it exists) dr=duis a vector in the direction of the tangent to the curve at …x1;x2;x3†.I fuis the arc length smeasured from some fixed point on the curve /C67, then dr=dsˆ^Tis a unit tangent vector to the curve /C67. The rate at which ^Tchanges with respect to sis a measure of the curvature of /C67and is given by d^T/ds. The direction of d^T/dsat any given point on /C67is normal to the curve at that point: ^T^Tˆ1,d…^T^T†=dsˆ0, from this we get^Td^T=dsˆ0, so they are normal to each other. If ^/C78is a unit vector in this normal direction (called the principal normal to the curve), then d^T=dsˆ/C20^/C78, and/C20is called the curvature of /C67at the specified point. The quantity /C26ˆ1=/C20is called the radius of curvature. In physics, we often study the motion of particles along curves, so the above results may be of value. In mechanics, the parameter uis time t, then dr=dtˆ/C118is the velocity of the particle which is tangent to the curve at the specific point. Now wecan write /C118ˆdr dtˆdr dsds dtˆ/C118^T where /C118is the magnitude of /C118, called the speed. Similarly, aˆd/C118=dtis the accel- eration of the particle. Motion in a plane Consider a particle Pmoving in a plane along a curve /C67(Fig. 1.12). Now rˆr^er, where ^eris a unit vector in the direction of r. Hence /C118ˆdr dtˆdr dt^er‡rd^er dt: 17MOTION IN A PLANE Figure 1.11. Parametric representation of a curve. Now d^er=dtis perpendicular to ^er. Also jd^er=dtjˆd=dt; we can easily verify this by di/C128erentiating ^erˆcos^e1‡sin^e2:Hence /C118ˆdr dtˆdr dt^er‡rd dt^e; ^eis a unit vector perpendicular to ^er. Di/C128erentiating again we obtain aˆd/C118 dtˆd2r dt2^er‡dr dtd^er dt‡dr dtd dt^e‡rd2 dt2^e‡rd dt^e ˆd2r dt2^er‡2dr dtd dt^e‡rd2 dt2^eÿrd dt2 ^er5d^e dtˆÿd dt^er : Thus aˆd2r dt2ÿrd dt2"# ^er‡1 rd dtr2d dt ^e: /C65 /C118ector treatment of classical orbit theor/C121 To illustrate the power and use of vector methods, we now employ them to work out the Keplerian orbits. We first prove Kepler’s second law which can be stated as: angular momentum is constant in a central force field. A central force is a force whose line of action passes through a single point or center and whose magnitude depends only on the distance from the center. Gravity and electrostatic forces are central forces. A general discussion on central force can be found in, for example, Chapter 6 of /C67lassical Mechanics , Tai L. Chow, John Wiley, New York, 1995. Di/C128erentiating the angular momentum Lˆrpwith respect to time, we obtain dL=dtˆdr=dtp‡rdp=dt: 18VECTOR AND TENSOR ANALYSIS Figure 1.12. Motion in a plane. The first vector product vanishes because pˆmdr=dtsodr=dtandpare parallel. The second vector product is simply rFby Newton’s second law, and hence vanishes for all forces directed along the position vector r, that is, for all central forces. Thus the angular momentum Lis a constant vector in central force motion. This implies that the position vector r, and therefore the entire orbit, lies in a fixed plane in three-dimensional space. This result is essentially Kepler’s second law, which is often stated in terms of the conservation of area velocity, jLj=2m. We now consider the inverse-square central force of gravitational and electro- statics. Newton’s second law then gives md/C118=dtˆÿ … k=r2†^n; …1:36† where ^nˆr=ris a unit vector in the r-direction, and kˆ/C71m 1m2for the gravita- tional force, and kˆ/C1131/C1132for the electrostatic force in cgs units. First we note that /C118ˆdr=dtˆdr=dt^n‡rd^n=dt: Then Lbecomes Lˆr…m/C118†ˆmr2‰^n…d^n=dt†Š: …1:37† Now consider d dt…/C118L†ˆd/C118 dtLˆÿk mr2…^nL†ˆÿk mr2‰^nmr2…^nd^n=dt†Š ˆÿk‰^n…d^n=dt^n†ÿ…d^n=dt†…^n^n†Š: Since ^n^nˆ1, it follows by di/C128erentiation that ^nd^n=dtˆ0. Thus we obtain d dt…/C118L†ˆkd^n=dt; integration gives /C118Lˆk^n‡/C67; …1:38† where /C67is a constant vector. It lies along, and fixes the position of, the major axis of the orbit as we shall see after we complete the derivation of the orbit. To find the orbit, we form the scalar quantity L2ˆL…rm/C118†ˆmr…/C118L†ˆmr…k‡Ccos†; …1:39† where is the angle measured from /C67(which we may take to be the x-axis) to r. Solving for r, we obtain rˆL2=km 1‡C=…kcos†ˆA 1‡/C34cos: …1:40† Eq. (1.40) is a conic section with one focus at the origin, where /C34represents the eccentricity of the conic section; depending on its values, the conic section may be 19A VECTOR TREATMENT OF CLASSICAL ORBIT THEORY a circle, an ellipse, a parabola, or a hyperbola. The eccentricity can be easily determined in terms of the constants of motion: /C34ˆC kˆ1 kj…/C118L†ÿk^nj ˆ1 k‰j/C118Lj2‡k2ÿ2k^n…/C118L†Š1=2 Now j/C118Lj2ˆ/C1182L2because /C118is perpendicular to L. Using Eq. (1.39), we obtain /C34ˆ1 k/C1182L2‡k2ÿ2kL2 mr"#1=2 ˆ1‡2L2 mk21 2m/C1182ÿk r"#1=2 ˆ1‡2L2/C69 mk2"#1=2 ; where Eis the constant energy of the system. /C86ector di/C128erentiation of a scalar field and the gradient Given a scalar field in a certain region of space given by a scalar function /C30…x1;x2;x3†that is defined and di/C128erentiable at each point with respect to the position coordinates …x1;x2;x3†, the total di/C128erential corresponding to an infini- tesimal change drˆ…dx1;dx2;dx3†is d/C30ˆ/C64/C30 /C64x1dx1‡/C64/C30 /C64x2dx2‡/C64/C30 /C64x3dx3: …1:41† We can express d/C30as a scalar product of two vectors: d/C30ˆ/C64/C30 /C64x1dx1‡/C64/C30 /C64x2dx2‡/C64/C30 /C64x3dx3ˆ/C114 /C30…†  dr; …1:42† where /C114/C30/C64/C30 /C64x1^e1‡/C64/C30 /C64x2^e2‡/C64/C30 /C64x3^e3 …1:43† is a vector field (or a vector point function). By this we mean to each pointrˆ…x 1;x2;x3†in space we associate a vector /C114/C30as specified by its three compo- nents ( /C64/C30=/C64 x1; /C64/C30=/C64 x2;/C64 /C30 = /C64 x3):/C114/C30is called the gradient of/C30and is often written as grad /C30. There is a simple geometric interpretation of /C114/C30. Note that /C30…x1;x2;x3†ˆc, where cis a constant, represents a surface. Let rˆx1^e1‡x2^e2‡x3^e3be the position vector to a point P…x1;x2;x3†on the surface. If we move along the surface to a nearby point /C81…r‡dr†, then drˆdx1^e1‡dx2^e2‡dx3^e3lies in the tangent plane to the surface at P. But as long as we move along the surface /C30has a constant value and d/C30ˆ0. Consequently from (1.41), dr/C114/C30ˆ0: …1:44† 20VECTOR AND TENSOR ANALYSIS Eq. (1.44) states that /C114/C30is perpendicular to drand therefore to the surface (Fig. 1.13). Let us return to d/C30ˆ… /C114 /C30†dr: The vector /C114/C30is fixed at any point P, so that d/C30, the change in /C30, will depend to a great extent on dr. Consequently d/C30will be a maximum when dris parallel to /C114/C30, since dr/C114/C30ˆjdrjj/C114/C30jcos, and cos is a maximum for ˆ0. Thus /C114/C30is in the direction of maximum increase of /C30…x1;x2;x3†. The component of /C114/C30in the direction of a unit vector ^uis given by /C114/C30^uand is called the directional deri- vative of /C30in the direction ^u. Physically, this is the rate of change of /C30at (x1;x2;x3†in the direction ^u. /C67onser/C118ati/C118e /C118ector field By definition, a vector field is said to be conservative if the line integral of the vector along any closed path vanishes. Thus, if Fis a conservative vector field (say, a conservative force field in mechanics), then I Fdsˆ0; …1:45† where dsis an element of the path. A (necessary and sucient) condition for F to be conservative is that Fcan be expressed as the gradient of a scalar, say /C30:Fˆÿgrad /C30: Zb aFdsˆÿZb agrad /C30dsˆÿZb ad/C30ˆ/C30…a†ÿ/C30…b†: it is obvious that the line integral depends solely on the value of the scalar /C30at the initial and final points, andH FdsˆÿH grad /C30dsˆ0. 21CONSERVATIVE VECTOR FIELD Figure 1.13. Gradient of a scalar. /C84he /C118ector di/C128erential operator /C114 We denoted the operation that changes a scalar field to a vector field in Eq. (1.43) by the symbol /C114(del or nabla): /C114/C64 /C64x1^e1‡/C64 /C64x2^e2‡/C64 /C64x3^e3; …1:46† which is called a gradient operator. We often write /C114/C30as grad /C30, and the vector field/C114/C30…r†is called the gradient of the scalar field /C30…r†. Notice that the operator /C114contains both partial di/C128erential operators and a direction: it is a vector di/C128er- ential operator. This important operator possesses properties analogous to those of ordinary vectors. It will help us in the future to keep in mind that /C114acts both as a di/C128erential operator and as a vector. /C86ector di/C128erentiation of a /C118ector field Vector di/C128erential operations on vector fields are more complicated because of the vector nature of both the operator and the field on which it operates. As we know there are two types of products involving two vectors, namely the scalar andvector products; vector di/C128erential operations on vector fields can also be sepa- rated into two types called the curl and the divergence. /C84he divergence of a vector If/C86…x 1;x2;x3†ˆ/C861^e1‡/C862^e2‡/C863^e3is a di/C128erentiable vector field (that is, it is defined and di/C128erentiable at each point ( x1;x2;x3) in a certain region of space), the divergence of /C86, written /C114/C86or div /C86, is defined by the scalar product /C114/C86ˆ/C64 /C64x1^e1‡/C64 /C64x2^e2‡/C64 /C64x3^e3 /C861^e1‡/C862^e2‡/C863^e3 …† ˆ/C64/C861 /C64x1‡/C64/C862 /C64x2‡/C64/C863 /C64x3: …1:47† The result is a scalar field. Note the analogy with /C65/C66ˆA1B1‡A2B2‡A3B3, but also note that /C114/C866ˆ/C86/C114(bear in mind that /C114is an operator). /C86/C114is a scalar di/C128erential operator: /C86/C114ˆ /C861/C64 /C64x1‡/C862/C64 /C64x2‡/C863/C64 /C64x3: What is the physical significance of the divergence/C63 Or why do we call the scalar product /C114/C86the divergence of /C86/C63 To answer these questions, we consider, as an example, the steady motion of a fluid of density /C26…x1;x2;x3†, and the velocity field is given by /C118…x1;x2;x3†ˆ/C1181…x1;x2;x3†e1‡/C1182…x1;x2;x3†e2‡/C1183…x1;x2;x3†e3.W e 22VECTOR AND TENSOR ANALYSIS now concentrate on the flow passing through a small parallelepiped AB/C67/C68E/C70/C71/C72 of dimensions dx1dx2dx3(Fig. 1.14). The x1andx3components of the velocity /C118 contribute nothing to the flow through the face AB/C67/C68 . The mass of fluid entering AB/C67/C68 per unit time is given by /C26/C1182dx1dx3and the amount leaving the face E/C70/C71/C72 per unit time is /C26/C1182‡/C64…/C26/C1182† /C64x2dx2 dx1dx3: So the loss of mass per unit time is ‰/C64…/C26/C1182†=/C64x2Šdx1dx2dx3. Adding the net rate of flow out all three pairs of surfaces of our parallelepiped, the total mass loss per unit time is /C64 /C64x1…/C26/C1181†‡/C64 /C64x2…/C26/C1182†‡/C64 /C64x3…/C26/C1183† dx1dx2dx3ˆ/C114… /C26/C118†dx1dx2dx3: So the mass loss per unit time per unit volume is /C114…/C26/C118†. Hence the name divergence. The divergence of any vector /C86is defined as /C114/C86. We now calculate /C114…f/C86†, where fis a scalar: /C114…f/C86†ˆ/C64 /C64x1…f/C861†‡/C64 /C64x2…f/C862†‡/C64 /C64x3…f/C863† ˆf/C64/C861 /C64x1‡/C64/C862 /C64x2‡/C64/C863 /C64x3 ‡/C861/C64f /C64x1‡/C862/C64f /C64x2‡/C863/C64f /C64x3 or /C114…f/C86†ˆf/C114/C86‡/C86/C114f: …1:48† It is easy to remember this result if we remember that /C114acts both as a di/C128erential operator and a vector. Thus, when operating on f/C86, we first keep ffixed and let /C114 23VECTOR DIFFERENTIATION OF A VECTOR FIELD Figure 1.14. Steady flow of a fluid. operate on /C86, and then we keep /C86fixed and let /C114operate on f…/C114 fis nonsense), and as /C114fand/C86are vectors we complete their multiplication by taking their dot product. A vector /C86is said to be solenoidal if its divergence is zero: /C114/C86ˆ0. /C84he operator /C1142/C44 the /C76aplacian The divergence of a vector field is defined by the scalar product of the operator /C114 with the vector field. What is the scalar product of /C114with itself /C63 /C1142ˆ/C114/C114ˆ/C64 /C64x1^e1‡/C64 /C64x2^e2‡/C64 /C64x3^e3 /C64 /C64x1^e1‡/C64 /C64x2^e2‡/C64 /C64x3^e3 ˆ/C642 /C64x2 1‡/C642 /C64x22‡/C642 /C64x23: This important quantity /C1142ˆ/C642 /C64x21‡/C642 /C64x22‡/C642 /C64x23…1:49† is a scalar di/C128erential operator which is called the Laplacian, after a French mathematician of the eighteenth century named Laplace. Now, what is the diver- gence of a gradient/C63 Since the Laplacian is a scalar di/C128erential operator, it does not change the vector character of the field on which it operates. Thus /C1142/C30…r†is a scalar field if/C30…r†is a scalar field, and /C1142‰/C114/C30…r†Šis a vector field because the gradient /C114/C30…r† is a vector field. The equation /C1142/C30ˆ0 is called Laplace’s equation. /C84he curl of a vector If/C86…x1;x2;x3†is a di/C128erentiable vector field, then the curl or rotation of /C86, written /C114/C86(or curl /C86or rot /C86), is defined by the vector product curl/C86ˆ/C114 /C86ˆ^e1 ^e2 ^e3 /C64 /C64x1/C64 /C64x2/C64 /C64x3 /C861/C862/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 ˆ^e 1/C64/C863 /C64x2ÿ/C64/C862 /C64x3 ‡^e2/C64/C861 /C64x3ÿ/C64/C863 /C64x1 ‡^e3/C64/C862 /C64x1ÿ/C64/C861 /C64x2 ˆX i;j;k/C34ijk^ei/C64/C86k /C64xj: …1:50† 24VECTOR AND TENSOR ANALYSIS The result is a vector field. In the expansion of the determinant the operators /C64=/C64ximust precede /C86i;P ijkstands forP iP jP k; and /C34ijkare the permutation symbols: an even permutation of ijkwill not change the value of the resulting permutation symbol, but an odd permutation gives an opposite sign. That is, /C34ijkˆ/C34jkiˆ/C34kijˆÿ/C34jikˆÿ/C34kjiˆÿ/C34ikj;and /C34ijkˆ0 if two or more indices are equal : A vector /C86is said to be irrotational if its curl is zero: /C114/C86…r†ˆ0. From this definition we see that the gradient of any scalar field /C30…r†is irrotational. The proof is simple: /C114… /C114 /C30†ˆ^e1 ^e2 ^e3 /C64 /C64x1/C64 /C64x2/C64 /C64x3 /C64 /C64x1/C64 /C64x2/C64 /C64x3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C30…x 1;x2;x3†ˆ0 …1:51† because there are two identical rows in the determinant. Or, in terms of the permutation symbols, we can write /C114… /C114 /C30†as /C114… /C114 /C30†ˆX ijk/C34ijk^ei/C64 /C64xj/C64 /C64xk/C30…x1;x2;x3†: Now /C34ijkis antisymmetric in j,k, but /C642=/C64xj/C64xkis symmetric, hence each term in the sum is always cancelled by another term: /C34ijk/C64 /C64xj/C64 /C64xk‡/C34ikj/C64 /C64xk/C64 /C64xjˆ0; and consequently /C114… /C114 /C30†ˆ0. Thus, for a conservative vector field F,w eh a v e curlFˆcurl (grad /C30†ˆ0. We learned above that a vector /C86is solenoidal (or divergence-free) if its diver- gence is zero. From this we see that the curl of any vector field /C86(r) must be solenoidal: /C114  …/C114  /C86†ˆX i/C64 /C64xi…/C114  /C86†iˆX i/C64 /C64xiX j;k/C34ijk/C64 /C64xj/C86k/C32! ˆ0; …1:52† because /C34ijkis antisymmetric in i,j. If/C30…r†is a scalar field and /C86(r) is a vector field, then /C114… /C30/C86†ˆ/C30…/C114  /C86†‡… /C114 /C30†/C86: …1:53† 25VECTOR DIFFERENTIATION OF A VECTOR FIELD We first write /C114… /C30/C86†ˆ^e1 ^e2 ^e3 /C64 /C64x1/C64 /C64x2/C64 /C64x3 /C30/C861/C30/C862/C30/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; then notice that /C64 /C64x1…/C30/C862†ˆ/C30/C64/C862 /C64x1‡/C64/C30 /C64x1/C862; so we can expand the determinant in the above equation as a sum of two deter- minants: /C114… /C30/C86†ˆ/C30^e1 ^e2 ^e3 /C64 /C64x1/C64 /C64x2/C64 /C64x3 /C861/C862/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12‡^e 1 ^e2 ^e3 /C64/C30 /C64x1/C64/C30 /C64x2/C64/C30 /C64x3 /C861/C862/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 ˆ/C30…/C114  /C86†‡… /C114 /C30†/C86: Alternatively, we can simplify the proof with the help of the permutation symbols /C34 ijk: /C114… /C30/C86†ˆX i;j;k/C34ij k^ei/C64 /C64xj…/C30/C86k† ˆ/C30X i;j;k/C34ij k^ei/C64/C86k /C64xj‡X i;j;k/C34ijk^ei/C64/C30 /C64xj/C86k ˆ/C30…/C114  /C86†‡… /C114 /C30†/C86: A vector field that has non-vanishing curl is called a vortex field, and the curl of the field vector is a measure of the vorticity of the vector field. The physical significance of the curl of a vector is not quite as transparent as that of the divergence. The following example from fluid flow will help us to develop a better feeling. Fig. 1.15 shows that as the component /C1182of the velocity /C118of the fluid increases with x3, the fluid curls about the x1-axis in a negative sense (rule of the right-hand screw), where /C64/C1182=/C64x3is considered positive. Similarly, a positive curling about the x1-axis would result from /C1183if/C64/C1183=/C64x2were positive. Therefore, the total x1component of the curl of /C118is ‰curl/C118Š1ˆ/C64/C1183=…/C64x2ÿ/C64/C1182=/C64x3; which is the same as the x1component of Eq. (1.50). 26VECTOR AND TENSOR ANALYSIS Formulas in/C118ol/C118ing /C114 We now list some important formulas involving the vector di/C128erential operator /C114, some of which are recapitulation. In these formulas, /C65and/C66are di/C128erentiable vector field functions, and fand gare di/C128erentiable scalar field functions of position …x1;x2;x3†: (1)/C114…f/C103†ˆf/C114/C103‡/C103/C114f; (2)/C114…f/C65†ˆf/C114/C65‡/C114f/C65; (3)/C114… f/C65†ˆf/C114/C65‡/C114f/C65; (4)/C114… /C114 f†ˆ0; (5)/C114… /C114 /C65†ˆ0; (6)/C114…/C65/C66† ˆ …/C114  /C65†/C66ÿ… /C114 /C66†/C65; (7)/C114… /C65/C66†ˆ… /C66/C114 †/C65ÿ/C66…/C114 /C65†‡/C65…/C114 /C66†ÿ…/C65/C114 †/C66; (8)/C114  …/C114  /C65† ˆ /C114…/C114  /C65†ÿ/C1142/C65; (9)/C114…/C65/C66†ˆ/C65… /C114 /C66†‡/C66… /C114 /C65†‡…/C65/C114 †/C66‡…/C66/C114 †/C65; (10)…/C65/C114 †rˆ/C65; (11)/C114rˆ3; (12)/C114rˆ0; (13)/C114…rÿ3r†ˆ0; (14) dFˆ…dr/C114 †F‡/C64F /C64tdt…Fa di/C128erentiable vector field quantity); (15) d’ˆdr/C114’‡/C64’ /C64tdt(’a di/C128erentiable scalar field quantity). /C79rthogonal cur/C118ilinear coordinates Up to this point all calculations have been performed in rectangular Cartesian coordinates. Many calculations in physics can be greatly simplified by using, instead of the familiar rectangular Cartesian coordinate system, another kind of 27FORMULAS INVOLVING /C114 Figure 1.15. Curl of a fluid flow. system which takes advantage of the relations of symmetry involved in the parti- cular problem under consideration. For example, if we are dealing with sphere, wewill find it expedient to describe the position of a point in sphere by the spherical coordinates ( r; ;/C30†. Spherical coordinates are a special case of the orthogonal curvilinear coordinate system. Let us now proceed to discuss these more generalcoordinate systems in order to obtain expressions for the gradient, divergence, curl, and Laplacian. Let the new coordinates u 1;u2;u3be defined by specifying the Cartesian coordinates ( x1;x2;x3) as functions of ( u1;u2;u3†: x1ˆf…u1;u2;u3†;x2ˆ/C103…u1;u2;u3†;x3ˆ/C104…u1;u2;u3†; …1:54† where f,g,hare assumed to be continuous, di/C128erentiable. A point P(Fig. 1.16) in space can then be defined not only by the rectangular coordinates ( x1;x2;x3) but also by curvilinear coordinates ( u1;u2;u3). Ifu2andu3are constant as u1varies, P(or its position vector r) describes a curve which we call the u1coordinate curve. Similarly, we can define the u2andu3coordi- nate curves through P. We adopt the convention that the new coordinate system is a right handed system, like the old one. In the new system drtakes the form: drˆ/C64r /C64u1du1‡/C64r /C64u2du2‡/C64r /C64u3du3: The vector /C64r=/C64u1is tangent to the u1coordinate curve at P.I f^u1is a unit vector atPin this direction, then ^u1ˆ/C64r=/C64u1=j/C64r=/C64u1j, so we can write /C64r=/C64u1ˆ/C1041^u1, where /C1041ˆj/C64r=/C64u1j. Similarly we can write /C64r=/C64u2ˆ/C1042^u2and /C64r=/C64u3ˆ/C1043^u3, where /C1042ˆj/C64r=/C64u2jand/C1043ˆj/C64r=/C64u3j, respectively. Then drcan be written drˆ/C1041du1^u1‡/C1042du2^u2‡/C1043du3^u3: …1:55† 28VECTOR AND TENSOR ANALYSIS Figure 1.16. Curvilinear coordinates. The quantities /C1041;/C1042;/C1043are sometimes called scale factors. The unit vectors ^u1,^u2, ^u3are in the direction of increasing u1;u2;u3, respectively. If^u1,^u2,^u3are mutually perpendicular at any point P, the curvilinear coordi- nates are called orthogonal. In such a case the element of arc length dsis given by ds2ˆdrdrˆ/C1042 1du21‡/C10422du22‡/C10423du23: …1:56† Along a u1curve, u2andu3are constants so that drˆ/C1041du1^u1. Then the di/C128erential of arc length ds1along u1atPis/C1041du1. Similarly the di/C128erential arc lengths along u2andu3atPareds2ˆ/C1042du2,ds3ˆ/C1043du3respectively. The volume of the parallelepiped is given by d/C86ˆj …/C1041du1^u1†…/C1042du2^u2†…/C1043du3^u3†j ˆ/C1041/C1042/C1043du1du2du3 since j^u1^u2^u3jˆ1. Alternatively d/C86can be written as d/C86ˆ/C64r /C64u1/C64r /C64u2/C64r /C64u3/C12/C12/C12/C12/C12/C12/C12/C12du 1du2du3ˆ/C64…x1;x2;x3† /C64…u1;u2;u3†/C12/C12/C12/C12/C12/C12/C12/C12du 1du2du3; …1:57† where /C74ˆ/C64…x1;x2;x3† /C64…u1;u2;u3†ˆ/C64x1 /C64u1/C64x1 /C64u2/C64x1 /C64u3 /C64x2 /C64u1/C64x2 /C64u2/C64x2 /C64u3 /C64x3 /C64u1/C64x3 /C64u2/C64x3 /C64u3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 is called the Jacobian of the transformation. We assume that the Jacobian /C746ˆ0 so that the transformation (1.54) is one to one in the neighborhood of a point. We are now ready to express the gradient, divergence, and curl in terms of u 1;u2,a n d u3.I f/C30is a scalar function of u1;u2,a n d u3, then the gradient takes the form /C114/C30ˆgrad /C30ˆ1 /C1041/C64/C30 /C64u1^u1‡1 /C1042/C64/C30 /C64u2^u2‡1 /C1043/C64/C30 /C64u3^u3: …1:58† To derive this, let /C114/C30ˆf1^u1‡f2^u2‡f3^u3; …1:59† where f1;f2;f3are to be determined. Since drˆ/C64r /C64u1du1‡/C64r /C64u2du2‡/C64r /C64u3du3 ˆ/C1041du1^u1‡/C1042du2^u2‡/C1043du3^u3; 29ORTHOGONAL CURVILINEAR COORDINATES we have d/C30ˆ/C114/C30drˆ/C1041f1du1‡/C1042f2du2‡/C1043f3du3: But d/C30ˆ/C64/C30 /C64u1du1‡/C64/C30 /C64u2du2‡/C64/C30 /C64u3du3; and on equating the two equations, we find fiˆ1 /C104i/C64/C30 /C64ui;iˆ1;2;3: Substituting these into Eq. (1.57), we obtain the result Eq. (1.58). From Eq. (1.58) we see that the operator /C114takes the form /C114ˆ^u1 /C1041/C64 /C64u1‡^u2 /C1042/C64 /C64u2‡^u3 /C1043/C64 /C64u3: …1:60† Because we will need them later, we now proceed to prove the following two relations: (a)j/C114uijˆ/C104ÿ1 i;iˆ1, 2, 3. (b)^u1ˆ/C1042/C1043/C114u2/C114u3with similar equations for ^u2and ^u3. (1.61) Proof: ( a) Let /C30ˆu1in Eq. (1.51), we then obtain /C114u1ˆ^u1=/C1041and so j/C114u1jˆj ^u1j/C104ÿ1 1ˆ/C104ÿ1 1;since j^u1jˆ1: Similarly by letting /C30ˆu2andu3, we obtain the relations for iˆ2 and 3. (b) From ( a) we have /C114u1ˆ^u1=/C1041;/C114u2ˆ^u2=/C1042;and /C114u3ˆ^u3=/C1043: Then /C114u2/C114u3ˆ^u2^u3 /C1042/C1043ˆ^u1 /C1042/C1043and ^u1ˆ/C1042/C1043/C114u2/C114u3: Similarly ^u2ˆ/C1043/C1041/C114u3/C114u1and ^u3ˆ/C1041/C1042/C114u1/C114u2: We are now ready to express the divergence in terms of curvilinear coordinates. If/C65ˆA1^u1‡A2^u2‡A3^u3is a vector function of orthogonal curvilinear coordi- nates u1,u2, and u3, the divergence will take the form /C114/C65ˆdiv/C65ˆ1 /C1041/C1042/C1043/C64 /C64u1…/C1042/C1043A1†‡/C64 /C64u2…/C1043/C1041A2†‡/C64 /C64u3…/C1041/C1042A3† :…1:62† To derive (1.62), we first write /C114/C65as /C114/C65ˆ/C114… A1^u1†‡/C114… A2^u2†‡/C114… A3^u3†; …1:63† 30VECTOR AND TENSOR ANALYSIS then, because ^u1ˆ/C1041/C1042/C114u2/C114u3, we express /C114…A1^u1)a s /C114…A1^u1†ˆ/C114… A1/C1042/C1043/C114u2/C114u3†…^u1ˆ/C1042/C1043/C114u2/C114u3† ˆ/C114 … A1/C1042/C1043†/C114u2/C114u3‡A1/C1042/C1043/C114… /C114 u2/C114u3†; where in the last step we have used the vector identity: /C114…/C30/C65†ˆ …/C114/C30†/C65‡/C30…/C114  /C65†.N o w /C114uiˆ^ui=/C104i;iˆ1, 2, 3, so /C114…A1^u1) can be rewritten as /C114…A1^u1†ˆ/C114 … A1/C1042/C1043†^u2 /C1042^u3 /C1043‡0ˆ/C114 … A1/C1042/C1043†^u1 /C1042/C1043: The gradient /C114…A1/C1042/C1043†is given by Eq. (1.58), and we have /C114…A1^u1†ˆ^u1 /C1041/C64 /C64u1…A1/C1042/C1043†‡^u2 /C1042/C64 /C64u2…A1/C1042/C1043†‡^u3 /C1043/C64 /C64u3…A1/C1042/C1043† ^u1 /C1042/C1043 ˆ1 /C1041/C1042/C1043/C64 /C64u1…A1/C1042/C1043†: Similarly, we have /C114…A2^u2†ˆ1 /C1041/C1042/C1043/C64 /C64u2…A2/C1043/C1041†;and /C114…A3^u3†ˆ1 /C1041/C1042/C1043/C64 /C64u3…A3/C1042/C1041†: Substituting these into Eq. (1.63), we obtain the result, Eq. (1.62). In the same manner we can derive a formula for curl /C65. We first write it as /C114/C65ˆ/C114… A1^u1‡A2^u2‡A3^u3† and then evaluate /C114Ai^ui. Now ^uiˆ/C104i/C114ui;iˆ1, 2, 3, and we express /C114… A1^u1†as /C114… A1^u1†ˆ/C114… A1/C1041/C114u1† ˆ/C114 … A1/C1041†/C114 u1‡A1/C1041/C114/C114 u1 ˆ/C114 … A1/C1041†^u1 /C1041‡0 ˆ^u1 /C1041/C64 /C64u1A1/C1041…† ‡^u2 /C1042/C64 /C64u2A2/C1042…† ‡^u3 /C1043/C64 /C64u3A3/C1043…† ^u1 /C1041 ˆ^u2 /C1043/C1041/C64 /C64u3A1/C1041…† ÿ^u3 /C1041/C1042/C64 /C64u2…A1/C1041†; 31ORTHOGONAL CURVILINEAR COORDINATES with similar expressions for /C114… A2^u2†and/C114… A3^u3†. Adding these together, we get /C114/C65in orthogonal curvilinear coordinates: /C114/C65ˆ^u1 /C1042/C1043/C64 /C64u2A3/C1043…† ÿ/C64 /C64u3A2/C1042…† ‡^u2 /C1043/C1041/C64 /C64u3A1/C1041…† ÿ/C64 /C64u1A3/C1043…† ‡^u3 /C1041/C1042/C64 /C64u1A2/C1042…† ÿ/C64 /C64u2…A1/C1041† : …1:64† This can be written in determinant form: /C114/C65ˆ1 /C1041/C1042/C1043/C1041^u1/C1042^u2/C1043^u3 /C64 /C64u1/C64 /C64u2/C64 /C64u3 A1/C1041A2/C1042A3/C1043/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: …1:65† We now express the Laplacian in orthogonal curvilinear coordinates. From Eqs. (1.58) and (1.62) we have /C114/C30ˆgrad /C30ˆ 1 /C1041/C64/C30 /C64u1^u1‡1 /C1042/C64/C30 /C64u2^u‡1 /C1043/C64/C30 /C64u3^u3; /C114/C65ˆdiv/C65ˆ1 /C1041/C1042/C1043/C64 /C64u1…/C1042/C1043A1†‡/C64 /C64u2…/C1043/C1041A2†‡/C64 /C64u3…/C1041/C1042A3† : If/C65ˆ/C114/C30, then Aiˆ…1=/C104i†/C64/C30=/C64 ui,iˆ1, 2, 3; and /C114/C65ˆ/C114/C114 /C30ˆ/C1142/C30 ˆ1 /C1041/C1042/C1043/C64 /C64u1/C1042/C1043 /C1041/C64/C30 /C64u1 ‡/C64 /C64u2/C1043/C1041 /C1042/C64/C30 /C64u2 ‡/C64 /C64u3/C1041/C1042 /C1043/C64/C30 /C64u3  :…1:66† /C83pecial orthogonal coordinate s/C121stems There are at least nine special orthogonal coordinates systems, the most common and useful ones are the cylindrical and spherical coordinates; we introduce these two coordinates in this section. /C67ylindrical coordinates …/C26; /C30;z† u1ˆ/C26;u2ˆ/C30;u3ˆz;and ^u1ˆe/C26;^u2ˆe/C30^u3ˆez: From Fig. 1.17 we see that x1ˆ/C26cos/C30;x2ˆ/C26sin/C30;x3ˆz 32VECTOR AND TENSOR ANALYSIS where /C260;0/C302;ÿ1 <z<1: The square of the element of arc length is given by ds2ˆ/C1042 1…d/C26†2‡/C10422…d/C30†2‡/C10423…dz†2: To find the scale factors /C104i, we notice that ds2ˆdrdrwhere rˆ/C26cos/C30e1‡/C26sin/C30e2‡ze3: Thus ds2ˆdrdrˆ…d/C26†2‡/C262…d/C30†2‡…dz†2: Equating the two ds2, we find the scale factors: /C1041ˆ/C104/C26ˆ1;/C1042ˆ/C104/C30ˆ/C26;/C1043ˆ/C104zˆ1: …1:67† From Eqs. (1.58), (1.62), (1.64), and (1.66) we find the gradient, divergence, curl, and Laplacian in cylindrical coordinates: /C114ˆ/C64 /C64/C26e/C26‡1 /C26/C64 /C64/C30e/C30‡/C64 /C64zez; …1:68† where ˆ…/C26; /C30;z†is a scalar function; /C114/C65ˆ1 /C26/C64 /C64/C26…/C26A/C26†‡/C64A/C30 /C64/C30‡/C64 /C64z…/C26Az† ; …1:69† 33SPECIAL ORTHOGONAL COORDINATE SYSTEMS Figure 1.17. Cylindrical coordinates. where /C65ˆA/C26e/C26‡A/C30e/C30‡Azez; /C114/C65ˆ1 /C26e/C26/C26e/C30ez /C64 /C64/C26/C64 /C64/C30/C64 /C64z A/C26/C26A/C30Az/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; …1:70† and /C114 2ˆ1 /C26/C64 /C64/C26/C26/C64 /C64/C26 ‡1 /C262/C642 /C64/C302‡/C642 /C64z2: …1:71† Spherical coordinates …r; ;/C30† u1ˆr;u2ˆ;u3ˆ/C30;^u1ˆer;^u2ˆe;^u3ˆe/C30 From Fig. 1.18 we see that x1ˆrsincos/C30;x2ˆrsinsin/C30;x3ˆrcos: Now ds2ˆ/C1042 1…dr†2‡/C10422…d†2‡/C10423…d/C30†2 but rˆrsincos/C30^e1‡rsinsin/C30^e2‡rcos^e3; 34VECTOR AND TENSOR ANALYSIS Figure 1.18. Spherical coordinates. so ds2ˆdrdrˆ…dr†2‡r2…d†2‡r2sin2…d/C30†2: Equating the two ds2, we find the scale factors: /C1041ˆ/C104rˆ1,/C1042ˆ/C104ˆr, /C1043ˆ/C104/C30ˆrsin. We then find, from Eqs. (1.58), (1.62), (1.64), and (1.66), the gradient, divergence, curl, and the Laplacian in spherical coordinates: /C114ˆ^er/C64 /C64r‡^e1 r/C64 /C64‡^e/C301 rsin/C64 /C64/C30; …1:72† /C114/C65ˆ1 r2sinsin/C64 /C64r…r2Ar†‡r/C64 /C64…sinA†‡r/C64A/C30 /C64/C30 ; …1:73† /C114/C65ˆ1 r2sin^err^ersin^e/C30 /C64 /C64r/C64 /C64/C64 /C64/C30 ArrArrsinA/C30/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; …1:74† /C114 2ˆ1 r2sinsin/C64 /C64rr2/C64 /C64r ‡/C64 /C64sin/C64 /C64 ‡1 sin/C642 /C64/C302"# : …1:75† /C86ector integration and integral theorems Having discussed vector di/C128erentiation, we now turn to a discussion of vector integration. After defining the concepts of line, surface, and volume integrals of vector fields, we then proceed to the important integral theorems of Gauss, Stokes, and Green. The integration of a vector, which is a function of a single scalar u, can proceed as ordinary scalar integration. Given a vector /C65…u†ˆA1…u†^e1‡A2…u†^e2‡A3…u†^e3; then Z /C65…u†duˆ^e1Z A1…u†du‡^e2Z A2…u†du‡^e3Z A3…u†du‡/C66; where /C66is a constant of integration, a constant vector. Now consider the integral of the scalar product of a vector /C65…x1;x2;x3) and drbetween the limit P1…x1;x2;x3) and P2…x1;x2;x3†: 35VECTOR INTEGRATION AND INTEGRAL THEOREMS ZP2 P1/C65drˆZP2 P1…A1^e1‡A2^e2‡A3^e3†…dx1^e1‡dx2^e2‡dx3^e3† ˆZP2 P1A1…x1;x2;x3†dx1‡ZP2 P1A2…x1;x2;x3†dx2 ‡ZP2 P1A3…x1;x2;x3†dx3: Each integral on the right hand side requires for its execution more than a knowl- edge of the limits. In fact, the three integrals on the right hand side are not completely defined because in the first integral, for example, we do not theknow value of x 2andx3inA1: I1ˆZP2 P1A1…x1;x2;x3†dx1: …1:76† What is needed is a statement such as x2ˆf…x1†;x3ˆ/C103…x1†… 1:77† that specifies x2,x3for each value of x1. The integrand now reduces to A1…x1;x2;x3†ˆA1…x1;f…x1†;/C103…x1†† ˆB1…x1†so that the integral I1becomes well defined. But its value depends on the constraints in Eq. (1.77). The con-straints specify paths on the x 1x2andx3x1planes connecting the starting point P1to the end point P2. The x1integration in (1.76) is carried out along these paths. It is a path-dependent integral and is called a line integral (or a path integral). It is very helpful to keep in mind that: when the number of integration variables is less than the number of variables in the integrand/C44 the integral is not yet completely de/C174ned and it is path/C45dependent . However, if the scalar product /C65dris equal to an exact di/C128erential, /C65drˆd’ˆ/C114’dr, the integration depends only upon the limits and is therefore path-independent: ZP2 P1/C65drˆZP2 P1d’ˆ’2ÿ’1: A vector field /C65which has above (path-independent) property is termed conser- vative. It is clear that the line integral above is zero along any close path, and the curl of a conservative vector field is zero …/C114  /C65ˆ/C114… /C114 ’†ˆ0†. A typical example of a conservative vector field in mechanics is a conservative force. The surface integral of a vector function /C65…x1;x2;x3†over the surface Sis an important quantity; it is defined to be Z S/C65da; 36VECTOR AND TENSOR ANALYSIS where the surface integral symbolR sstands for a double integral over a certain surface S, and dais an element of area of the surface (Fig. 1.19), a vector quantity. We attribute to daa magnitude daand also a direction corresponding the normal, ^n, to the surface at the point in question, thus daˆ^nda: The normal ^nto a surface may be taken to lie in either of two possible directions. But if dais part of a closed surface, the sign of ^nrelative to dais so chosen that it points outward away from the interior. In rectangular coordinates we may write daˆ^e1da1‡^e2da2‡^e3da3ˆ^e1dx2dx3‡^e2dx3dx1‡^e3dx1dx2: If a surface integral is to be evaluated over a closed surface S, the integral is written as I S/C65da: Note that this is di/C128erent from a closed-path line integral. When the path of integration is closed, the line integral is write it as I ÿ^/C65ds; where ÿspecifies the closed path, and dsis an element of length along the given path. By convention, dsis taken positive along the direction in which the path is traversed. Here we are only considering simple closed curves. A simple closed curve does not intersect itself anywhere. Gauss’ theorem /C40the divergence theorem/C41 This theorem relates the surface integral of a given vector function and the volumeintegral of the divergence of that vector. It was introduced by Joseph Louis Lagrange and was first used in the modern sense by George Green. Gauss’ 37VECTOR INTEGRATION AND INTEGRAL THEOREMS Figure 1.19. Surface integral over a surface S. name is associated with this theorem because of his extensive work on general problems of double and triple integrals. If a continuous, di/C128erentiable vector field /C65is defined in a simply connected region of volume /C86bounded by a closed surface S, then the theorem states that Z /C86/C114/C65d/C86ˆI S/C65da; …1:78† where d/C86ˆdx1dx2dx3. A simple connected region /C86has the property that every simple closed curve within it can be continuously shrunk to a point withoutleaving the region. To prove this, we first write Z /C86/C114/C65d/C86ˆZ /C86X3 iˆ1/C64Ai /C64xid/C86; then integrate the right hand side with respect to x1while keeping x2x3constant, thus summing up the contribution from a rod of cross section dx2dx3(Fig. 1.20). The rod intersects the surface Sat the points Pand/C81and thus defines two elements of area daPandda/C81: Z /C86/C64A1 /C64x1d/C86ˆI Sdx2dx3Z/C81 P/C64A1 /C64x1dx1ˆI Sdx2dx3Z/C81 PdA1; where we have used the relation dA1ˆ…/C64A1=/C64x1†dx1along the rod. The last integration on the right hand side can be performed at once and we have Z /C86/C64A1 /C64x1d/C86ˆI S‰A1…/C81†ÿA1…P†Šdx2dx3; where A1…/C81†denotes the value of A1evaluated at the coordinates of the point /C81, and similarly for A1…P†. The component of the surface element dawhich lies in the x1-direction is da1ˆdx2dx3at the point /C81, and da1ˆÿdx2dx3at the point P. The minus sign 38VECTOR AND TENSOR ANALYSIS Figure 1.20. A square tube of cross section dx2dx3. arises since the x1component of daatPis in the direction of negative x1.W ec a n now rewrite the above integral as Z /C86/C64A1 /C64x1d/C86ˆZ S/C81A1…/C81†da1‡Z SPA1…P†da1; where S/C81denotes that portion of the surface for which the x1component of the outward normal to the surface element da1is in the positive x1-direction, and SP denotes that portion of the surface for which da1is in the negative direction. The two surface integrals then combine to yield the surface integral over the entire surface S(if the surface is suciently concave, there may be several such as right hand and left hand portions of the surfaces): Z /C86/C64A1 /C64x1d/C86ˆI SA1da1: Similarly we can evaluate the x2andx3components. Summing all these together, we have Gauss’ theorem: Z /C86X i/C64Ai /C64xid/C86ˆI SX iAidai orZ /C86/C114/C65d/C86ˆI S/C65da: We have proved Gauss’ theorem for a simply connected region (a volume bounded by a single surface), but we can extend the proof to a multiply connectedregion (a region bounded by several surfaces, such as a hollow ball). For inter- ested readers, we recommend the book Electromagnetic /C70ields , Roald K. Wangsness, John Wiley, New York, 1986. /C67ontinuity equation Consider a fluid of density /C26…r†which moves with velocity /C118(r) in a certain region. If there are no sources or sinks, the following continuity equation must be satis-fied: /C64/C26…r†=/C64t‡/C114 /C106…r†ˆ0; …1:79† where /C106is the current /C106…r†ˆ/C26…r†/C118…r†… 1:79a† and Eq. (1.79) is called the continuity equation for a conserved current. To derive this important equation, let us consider an arbitrary surface Senclos- ing a volume /C86of the fluid. At any time the mass of fluid within /C86isMˆR /C86/C26d/C86 and the time rate of mass increase (due to mass flowing into /C86)i s /C64M /C64tˆ/C64 /C64tZ /C86/C26d/C86ˆZ /C86/C64/C26 /C64td/C86; 39VECTOR INTEGRATION AND INTEGRAL THEOREMS while the mass of fluid leaving /C86per unit time isZ S/C26/C118^ndsˆZ /C86/C114…/C26/C118†d/C86; where Gauss’ theorem is used in changing the surface integral to volume integral. Since there is neither a source nor a sink, mass conservation requires an exact balance between these e/C128ects:Z /C86/C64/C26 /C64td/C86ˆÿZ /C86/C114…/C26/C118†d/C86;orZ /C86/C64/C26 /C64t‡/C114… /C26/C118† d/C86ˆ0: Also since /C86is arbitrary, mass conservation requires that the continuity equation /C64/C26 /C64t‡/C114… /C26/C118†ˆ/C64/C26 /C64t/C114/C106ˆ0 must be satisfied everywhere in the region. Sto/C107es’ theorem This theorem relates the line integral of a vector function and the surface integralof the curl of that vector. It was first discovered by Lord Kelvin in 1850 and rediscovered by George Gabriel Stokes four years later. If a continuous, di/C128erentiable vector field /C65is defined a three-dimensional region /C86, and Sis a regular open surface embedded in /C86bounded by a simple closed curve ÿ, the theorem states thatZ S/C114/C65daˆI ÿ/C65dl; …1:80† where the line integral is to be taken completely around the curve ÿanddlis an element of line (Fig. 1.21). 40VECTOR AND TENSOR ANALYSIS Figure 1.21. Relation between daanddlin defining curl. The surface S, bounded by a simple closed curve, is an open surface; and the normal to an open surface can point in two opposite directions. We adopt the usual convention, namely the right hand rule: when the fingers of the right hand follow the direction of dl, the thumb points in the dadirection, as shown in Fig. 1.21. Note that Eq. (1.80) does not specify the shape of the surface Sother than that it be bounded by ÿ; thus there are many possibilities in choosing the surface. But Stokes’ theorem enables us to reduce the evaluation of surface integrals which depend upon the shape of the surface to the calculation of a line integral which depends only on the values of /C65along the common perimeter. To prove the theorem, we first expand the left hand side of Eq. (1.80); with the aid of Eq. (1.50), it becomes Z S/C114/C65daˆZ S/C64A1 /C64x3da2ÿ/C64A1 /C64x2da3 ‡Z S/C64A2 /C64x1da3ÿ/C64A2 /C64x3da1 ‡Z S/C64A3 /C64x2da1ÿ/C64A3 /C64x1da2 ; …1:81† where we have grouped the terms by components of /C65. We next subdivide the surface Sinto a large number of small strips, and integrate the first integral on the right hand side of Eq. (1.81), denoted by I1, over one such a strip of width dx1, which is parallel to the x2x3plane and a distance x1from it, as shown in Fig. 1.21. Then, by integrating over x1, we sum up the contributions from all of the strips. Fig. 1.21 also shows the projections of the strip on the x1x3andx1x2planes that will help us to visualize the orientation of the surface. The element area dais shown at an intermediate stage of the integration, when the direction angles havevalues such that and /C13are less than 90 8and /C12is greater than 90 8. Thus, da 2ˆÿdx1dx3andda3ˆdx1dx2and we can write I1ˆÿZ stripsdx1Z/C81 P/C64A1 /C64x2dx2‡/C64A1 /C64x3dx3 : …1:82† Note that dx2anddx3in the parentheses are not independent because x2andx3 are related by the equation for the surface Sand the value of x1involved. Since the second integral in Eq. (1.82) is being evaluated on the strip from Pto/C81for which x1ˆconst., dx1ˆ0 and we can add …/C64A1=/C64x1†dx1ˆ0 to the integrand to make it dA1: /C64A1 /C64x1dx1‡/C64A1 /C64x2dx2‡/C64A1 /C64x3dx3ˆdA1: And Eq. (1.82) becomes I1ˆÿZ stripsdx1Z/C81 PdA1ˆZ stripsA1…P†ÿA1…/C81† ‰Š dx1: 41VECTOR INTEGRATION AND INTEGRAL THEOREMS Next we consider the line integral of /C65around the lines bounding each of the small strips. If we trace each one of these lines in the same sense as we trace the path ÿ, then we will trace all of the interior lines twice (once in each direction) and all of the contributions to the line integral from the interior lines will cancel,leaving only the result from the boundary line ÿ. Thus, the sum of all of the line integrals around the small strips will equal the line integral ÿofA 1: Z S/C64A1 /C64x3da2ÿ/C64A1 /C64x2da3 ˆI ÿA1d/C1081: …1:83† Similarly, the last two integrals of Eq. (1.81) can be shown to have the respectivevalues I ÿA2d/C1082andI ÿA3d/C1083: Substituting these results and Eq. (1.83) into Eq. (1.81) we obtain Stokes’ theorem: Z S/C114/C65daˆI ÿ…A1d/C1081‡A2d/C1082‡A3d/C1083†ˆI ÿ/C65dl: Stokes’ theorem in Eq. (1.80) is valid whether or not the closed curve ÿlies in a plane, because in general the surface Sis not a planar surface. Stokes’ theorem holds for any surface bounded by ÿ. In fluid dynamics, the curl of the velocity field /C118…r†is called its vorticity (for example, the whirls that one creates in a cup of co/C128ee on stirring it). If the velocityfield is derivable from a potential /C118…r†ˆÿ /C114 /C30…r† it must be irrotational (see Eq. (1.51)). For this reason, an irrotational flow is alsocalled a potential flow, which describes a steady flow of the fluid, free of vortices and eddies. One of Maxwell’s equations of electromagnetism (Ampe /C193re’s law) states that /C114/C66ˆ/C22 0/C106; where /C66is the magnetic induction, /C106is the current density (per unit area), and /C220is the permeability of free space. From this equation, current densities may bevisualized as vortices of /C66. Applying Stokes’ theorem, we can rewrite Ampe /C193re’s law as I ÿ/C66drˆ/C220Z S/C106daˆ/C220I; it states that the circulation of the magnetic induction is proportional to the total current Ipassing through the surface Senclosed by ÿ. 42VECTOR AND TENSOR ANALYSIS Green’s theorem Green’s theorem is an important corollary of the divergence theorem, and it has many applications in many branches of physics. Recall that the divergence theorem Eq. (1.78) states that Z /C86/C114/C65d/C86ˆI S/C65da: Let/C65ˆ/C32/C66, where /C32is a scalar function and /C66a vector function, then /C114/C65 becomes /C114/C65ˆ/C114… /C32/C66†ˆ/C32/C114/C66‡/C66/C114/C32: Substituting these into the divergence theorem, we have I S/C32/C66daˆZ /C86…/C32/C114/C66‡/C66/C114/C32†d/C86: …1:84† If/C66represents an irrotational vector field, we can express it as a gradient of a scalar function, say, ’: /C66/C114’: Then Eq. (1.84) becomes I S/C32/C66daˆZ /C86‰/C32/C114… /C114 ’†‡… /C114 ’†… /C114 /C32†Šd/C86: …1:85† Now /C66daˆ… /C114 ’†^nda: The quantity …/C114’†^nrepresents the rate of change of /C30in the direction of the outward normal; it is called the normal derivative and is written as …/C114’†^n/C64’=/C64 n: Substituting this and the identity /C114… /C114 ’†ˆ/C1142’into Eq. (1.85), we have I S/C32/C64’ /C64ndaˆZ /C86‰/C32/C1142’‡/C114’/C114/C32Šd/C86: …1:86† Eq. (1.86) is known as Green’s theorem in the first form. Now let us interchange ’and/C32, then Eq. (1.86) becomes I S’/C64/C32 /C64ndaˆZ /C86‰’/C1142/C32‡/C114’/C114/C32Šd/C86: Subtracting this from Eq. (1.85): I S/C32/C64’ /C64nÿ’/C64/C32 /C64n daˆZ /C86/C32/C1142’ÿ’/C1142/C32ÿ d/C86: …1:87† 43VECTOR INTEGRATION AND INTEGRAL THEOREMS This important result is known as the second form of Green’s theorem, and has many applications. Green’s theorem in the plane Consider the two-dimensional vector field /C65ˆM…x1;x2†^e1‡/C78…x1;x2†^e2. From Stokes’ theorem I ÿ/C65drˆZ S/C114/C65daˆZ S/C64/C78 /C64x1ÿ/C64M /C64x2 dx1dx2; …1:88† which is often called Green’s theorem in the plane. SinceH ÿ/C65drˆH ÿ…Mdx 1‡/C78dx 2†, Green’s theorem in the plane can be writ- ten as I ÿMdx 1‡/C78dx 2ˆZ S/C64/C78 /C64x1ÿ/C64M /C64x2 dx1dx2: …1:88a† As an illustrative example, let us apply Green’s theorem in the plane to show that the area bounded by a simple closed curve ÿis given by 1 2I ÿx1dx2ÿx2dx1: Into Green’s theorem in the plane, let us put Mˆÿx2;/C78ˆx1, giving I ÿx1dx2ÿx2dx1ˆZ S/C64 /C64x1x1ÿ/C64 /C64x2…ÿx2† dx1dx2ˆ2Z Sdx1dx2ˆ2A; where Ais the required area. Thus Aˆ1 2H ÿx1dx2ÿx2dx1. /C72elmholt/C122/C39s theorem The divergence and curl of a vector field play very important roles in physics. We learned in previous sections that a divergence-free field is solenoidal and a curl- free field is irrotational. We may classify vector fields in accordance with their being solenoidal and/or irrotational. A vector field /C86is: (1) Solenoidal and irrotational if /C114/C86ˆ0a n d /C114/C86ˆ0. A static electric field in a charge-free region is a good example. (2) Solenoidal if /C114/C86ˆ0 but /C114/C866ˆ0. A steady magnetic field in a current- carrying conductor meets these conditions. (3) Irrotational if /C114/C86ˆ0 but /C114/C86ˆ0. A static electric field in a charged region is an irrotational field. The most general vector field, such as an electric field in a charged medium with a time-varying magnetic field, is neither solenoidal nor irrotational, but can be 44VECTOR AND TENSOR ANALYSIS considered as the sum of a solenoidal field and an irrotational field. This is made clear by Helmholtz’s theorem, which can be stated as (C. W. Wong: Introduction to Mathematical Physics , Oxford University Press, Oxford 1991; p. 53): A vector field is uniquely determined by its divergence and curl in a region of space, and its normal component over the boundary of the region. In particular, if both divergence and curl arespecified everywhere and if they both disappear at infinity suciently rapidly, then the vector field can be written as a unique sum of an irrotational part and a solenoidal part. In other words, we may write /C86…r†ˆÿ /C114 /C30…r†‡/C114 /C65…r†; …1:89† where ÿ/C114/C30is the irrotational part and /C114/C65is the solenoidal part, and /C30(r) and /C65…r†are called the scalar and the vector potential, respectively, of /C86…r). If both /C65 and/C30can be determined, the theorem is verified. How, then, can we determine /C65 and/C30/C63 If the vector field /C86…r†is such that /C114/C86…r†ˆ/C26;and /C114/C86…r†ˆ/C118; then we have /C114/C86…r†ˆ/C26ˆÿ /C114… /C114 /C30†‡/C114… /C114 /C65† or /C114 2/C30ˆÿ/C26; which is known as Poisson’s equation. Next, we have /C114/C86…r†ˆ/C118ˆ /C114  ‰ÿ/C114 /C30‡/C114 /C65…r†Š or /C1142/C65ˆ/C118; or in component, we have /C1142Aiˆ/C118i;iˆ1;2;3 where these are also Poisson’s equations. Thus, both /C65and/C30can be determined by solving Poisson’s equations. /C83ome useful integral relations These relations are closely related to the general integral theorems that we have proved in preceding sections. (1) The line integral along a curve /C67between two points aandbis given by Zb a/C114/C30…†  dlˆ/C30…b†ÿ/C30…a†: …1:90† 45SOME USEFUL INTEGRAL RELATIONS Proof: Zb a/C114/C30…†  dlˆZb a/C64/C30 /C64x^i‡/C64/C30 /C64y^j‡/C64/C30 /C64z^k …dx^i‡dy^j‡dz^k† ˆZb a/C64/C30 /C64xdx‡/C64/C30 /C64ydy‡/C64/C30 /C64zdz ˆZb a/C64/C30 /C64xdx dt‡/C64/C30 /C64ydy dt‡/C64/C30 /C64zdz dt dt ˆZb ad/C30 dt dtˆ/C30…b†ÿ/C30…a†: …2†I S/C64’ /C64ndaˆZ /C86/C1142’d/C86: …1:91† Proof: Set /C32ˆ1 in Eq. (1.87), then /C64/C32=/C64 nˆ0ˆ/C1142/C32and Eq. (1.87) reduces to Eq. (1.91). …3†Z /C86/C114’d/C86ˆI S’^nda: …1:92† Proof: In Gauss’ theorem (1.78), let /C65ˆ’/C67, where /C67is constant vector. Then we have Z /C86/C114…’/C67†d/C86ˆZ S’/C67^nda: Since /C114…’/C67†ˆ/C114 ’/C67ˆ/C67/C114’and ’/C67^nˆ/C67…’^n†; we have Z /C86/C67/C114’d/C86ˆZ S/C67…’^n†da: Taking /C67outside the integrals, /C67Z /C86/C114’d/C86ˆ/C67Z S…’^n†da and since /C67is an arbitrary constant vector, we have Z /C86/C114’d/C86ˆI S’^nda: …4†Z /C86/C114/C66d/C86ˆZ S^n/C66da …1:93† 46VECTOR AND TENSOR ANALYSIS Proof: In Gauss’ theorem (1.78), let /C65ˆ/C66/C67where /C67is a constant vector. We then have Z /C86/C114…/C66/C67†d/C86ˆZ S…/C66/C67†^nda: Since /C114…/C66/C67†ˆ/C67… /C114 /C66†and …/C66/C67†^nˆ/C66…/C67^n†ˆ… /C67^n†/C66ˆ /C67…^n/C66†; Z /C86/C67… /C114 /C66†d/C86ˆZ S/C67…^n/C66†da: Taking /C67outside the integrals /C67Z /C86…/C114  /C66†d/C86ˆ/C67Z S…^n/C66†da and since /C67is an arbitrary constant vector, we have Z /C86/C114/C66d/C86ˆZ S^n/C66da: /C84ensor anal/C121sis Tensors are a natural generalization of vectors. The beginnings of tensor analysis can be traced back more than a century to Gauss’ works on curved surfaces. Today tensor analysis finds applications in theoretical physics (for example, gen- eral theory of relativity, mechanics, and electromagnetic theory) and to certainareas of engineering (for example, aerodynamics and fluid mechanics). The gen- eral theory of relativity uses tensor calculus of curved space-time, and engineers mainly use tensor calculus of Euclidean space. Only general tensors are considered in this section. The general definition of a tensor is given, followed by a concise discussion of tensor algebra and tensor calculus (covariant di/C128eren- tiation). Tensors are defined by means of their properties of transformation under coordinate transformation. Let us consider the transformation from one coordi- nate system …x 1;x2;...;x/C78†to another …x01;x02;...;x0/C78†in an /C78-dimensional space /C86/C78. Note that in writing x/C22, the index /C22is a superscript and should not be mistaken for an exponent. In three-dimensional space we use subscripts. Wenow use superscripts in order that we may maintain a ‘balancing’ of the indices in all the general equations. The meaning of ‘balancing’ will become clear a little later. When we transform the coordinates, their di/C128erentials transform according to the relation dx /C22ˆ/C64x/C22 /C64x0/C23dx0/C23: …1:94† 47TENSOR ANALYSIS Here we have used Einstein’s summation convention: repeated indexes which appear once in the lower and once in the upper position are automaticallysummed over. Thus, X /C78 /C22ˆ1A/C22A/C22ˆA/C22A/C22: It is important to remember that indexes repeated in the lower part or upper part alone are not summed over. An index which is repeated and over which summa- tion is implied is called a dummy index. Clearly, a dummy index can be replaced by any other index that does not appear in the same term. /C67ontra/C118ariant and co/C118ariant /C118ectors A set of /C78quantities A/C22…/C22ˆ1;2;...;/C78†which, under a coordinate change, transform like the coordinate di/C128erentials, are called the components of a contra- variant vector or a contravariant tensor of the first rank or first order: A/C22ˆ/C64x/C22 /C64x0/C23A0/C23: …1:95† This relation can easily be inverted to express A0/C23in terms of A/C22. We shall leave this as homework for the reader (Problem 1.32). If/C78quantities A/C22…/C22ˆ1;2;...;/C78†in a coordinate system …x1;x2;...;x/C78†are related to /C78other quantities A0 /C23…/C23ˆ1;2;...;/C78†in another coordinate system …x01;x02;...;x0/C78†by the transformation equations A/C22ˆ/C64x0/C23 /C64x/C22A/C23 …1:96† they are called components of a covariant vector or covariant tensor of the firstrank or first order. One can show easily that velocity and acceleration are contravariant vectors and that the gradient of a scalar field is a covariant vector (Problem 1.33). Instead of speaking of a tensor whose components are A /C22orA/C22we shall simply refer to the tensor A/C22orA/C22. /C84ensors of second ran/C107 From two contravariant vectors A/C22andB/C23we may form the /C782quantities A/C22B/C23. This is known as the outer product of tensors. These /C782quantities form the components of a contravariant tensor of the second rank: any aggregate of /C782 quantities T/C22/C23which, under a coordinate change, transform like the product of 48VECTOR AND TENSOR ANALYSIS two contravariant vectors T/C22/C23ˆ/C64x/C22 /C64x0 /C64x/C23 /C64x0/C12T0 /C12; …1:97† is a contravariant tensor of rank two. We may also form a covariant tensor of rank two from two covariant vectors, which transforms according to the formula T/C22/C23ˆ/C64x0 /C64x/C22/C64x0/C12 /C64x/C23T0 /C12: …1:98† Similarly, we can form a mixed tensor T/C22 /C23of order two that transforms as follows: T/C22 /C23ˆ/C64x/C22 /C64x0 /C64x0/C12 /C64x/C23T0 /C12: …1:99† We may continue this process and multiply more than two vectors together, taking care that their indexes are all di/C128erent. In this way we can construct tensors of higher rank. The total number of free indexes of a tensor is its rank (or order). In a Cartesian coordinate system, the distinction between the contravariant and the covariant tensors vanishes. This can be illustrated with the velocity andgradient vectors. Velocity and acceleration are contravariant vectors, they are represented in terms of components in the directions of coordinate increase; the gradient vector is a covariant vector and it is represented in terms of components in the directions orthogonal to the constant coordinate surfaces. In a Cartesian coordinate system, the coordinate direction x /C22coincides with the direction ortho- gonal to the constant- x/C22surface, hence the distinction between the covariant and the contravariant vectors vanishes. In fact, this is the essential di/C128erence between contravariant and covariant tensors: a covariant tensor is represented by com- ponents in directions orthogonal to like constant coordinate surface, and acontravariant tensor is represented by components in the directions of coordinate increase. If two tensors have the same contravariant rank and the same covariant rank, we say that they are of the same type. /C66asic operations /C119ith tensors (1) Equality: Two tensors are said to be equal if and only if they have the same covariant rank and the same contravariant rank, and every component of one is equal to the corresponding component of the other: A /C12 /C22ˆB /C12 /C22: 49BASIC OPERATIONS WITH TENSORS (2) Addition (subtraction): The sum (di/C128erence) of two or more tensors of the same type and rank is also a tensor of the same type and rank. Addition of tensors is commutative and associative. (3) Outer product of tensors: The product of two tensors is a tensor whose rank is the sum of the ranks of the given two tensors. This product involves ordinary multiplication of the components of the tensor and it is called the outer product. For example, A/C22/C23 B/C12 ˆC/C22/C23 /C12is the outer product ofA/C22/C23 andB/C12 . (4) Contraction: If a covariant and a contravariant index of a mixed tensor are set equal, a summation over the equal indices is to be taken according to thesummation convention. The resulting tensor is a tensor of rank two less thanthat of the original tensor. This process is called contraction. For example, if we start with a fourth-order tensor T /C22 /C23/C26, one way of contracting it is to set ˆ/C26, which gives the second rank tensor T/C22 /C23/C26/C26. We could contract it again to get the scalar T/C22 /C22/C26/C26. (5) Inner product of tensors: The inner product of two tensors is produced by contracting the outer product of the tensors. For example, given two tensors A /C12 andB/C22 /C23, the outer product is A /C12 B/C22 /C23. Setting ˆ/C22, we obtain the inner product A /C12 /C22B/C22 /C23. (6) Symmetric and antisymmetric tensors: A tensor is called symmetric with respect to two contravariant or two covariant indices if its componentsremain unchanged upon interchange of the indices: A /C12ˆA/C12 ;A /C12ˆA/C12 : A tensor is called anti-symmetric with respect to two contravariant or two covariant indices if its components change sign upon interchange of the indices: A /C12ˆÿA/C12 ;A /C12ˆÿA/C12 : Symmetry and anti-symmetry can be defined only for similar indices, not when one index is up and the other is down. /C81uotient la/C119 A quantity /C81 ... /C22...with various up and down indexes may or may not be a tensor. We can test whether it is a tensor or not by using the quotient law, which can be stated as follows: Suppose it is not known whether a quantity Xis a tensor or not. If an inner product of Xwith an arbitrary tensor is a tensor, then Xis also a tensor. 50VECTOR AND TENSOR ANALYSIS As an example, let XˆP/C22/C23;Abe an arbitrary contravariant vector, and AP/C22/C23 be a tensor, say /C81/C22/C23:AP/C22/C23ˆ/C81/C22/C23, then AP/C22/C23ˆ/C64x0 /C64x/C22/C64x0/C12 /C64x/C23A0/C13P0 /C13 /C12: But A0/C13ˆ/C64x0/C13 /C64xA and so AP/C22/C23ˆ/C64x0 /C64x/C22/C64x0/C12 /C64x/C23/C64x0/C13 /C64xA0P0 /C13 /C12: This equation must hold for all values of A, hence we have, after canceling the arbitrary A, P/C22/C23ˆ/C64x0 /C64x/C22/C64x0/C12 /C64x/C23/C64x0/C13 /C64xP0 /C13 /C12; which shows that P/C22/C23is a tensor (contravariant tensor of rank 3). /C84he line element and metric tensor So far covariant and contravariant tensors have nothing to do each other except that their product is an invariant: A0 /C22B0/C22ˆ/C64x /C64x0/C22/C64x0/C22 /C64x/C12A A/C12ˆ/C64x /C64x/C12A A/C12ˆ /C12A A/C12ˆA A : A space in which covariant and contravariant tensors exist separately is called ane. Physical quantities are independent of the particular choice of the modeof description (that is, independent of the possible choice of contravariance orcovariance). Such a space is called a metric space. In a metric space, contravariant and covariant tensors can be converted into each other with the help of the metric tensor /C103 /C22/C23. That is, in metric spaces there exists the concept of a tensor that may be described by covariant indices, or by contravariant indices. These two descrip-tions are now equivalent. To introduce the metric tensor /C103 /C22/C23, let us consider the line element in /C86/C78.I n rectangular coordinates the line element (the di/C128erential of arc length) dsis given by ds2ˆdx2‡dy2‡dz2ˆ…dx1†2‡…dx2†2‡…dx3†2; there are no cross terms dxidxj. In curvilinear coordinates ds2cannot be repre- sented as a sum of squares of the coordinate di/C128erentials. As an example, in spherical coordinates we have ds2ˆdr2‡r2d2‡r2sin2d/C302 which can be in a quadratic form, with x1ˆr;x2ˆ;x3ˆ/C30. 51THE LINE ELEMENT AND METRIC TENSOR A generalization to /C86/C78is immediate. We define the line element dsin/C86/C78to be given by the following quadratic form, called the metric form, or metric ds2ˆX3 /C22ˆ1X3 /C23ˆ1/C103/C22/C23dx/C22dx/C23ˆ/C103/C22/C23dx/C22dx/C23: …1:100† For the special cases of rectangular coordinates and spherical coordinates, we have ~/C103ˆ…/C103/C22/C23†ˆ100 010 0010 B@1 CA; ~/C103ˆ…/C103/C22/C23†ˆ10 0 0r20 00 r2sin20 B@1 CA: …1:101† In an /C78-dimensional orthogonal coordinate system /C103/C22/C23ˆ0 for /C226ˆ/C23. And in a Cartesian coordinate system /C103/C22/C22ˆ1a n d /C103/C22/C23ˆ0 for /C226ˆ/C23. In the general case of Riemannian space, the /C103/C22/C23are functions of the coordinates x/C22…/C22ˆ1;2;...;/C78†. Since the inner product of /C103/C22/C23and the contravariant tensor dx/C22dx/C23is a scalar (ds2, the square of line element), then according to the quotient law /C103/C22/C23is a covariant tensor. This can be demonstrated directly: ds2ˆ/C103 /C12dx dx/C12ˆ/C1030 /C12dx0 dx0/C12: Now dx0 ˆ…/C64x0 =/C64x/C22†dx/C22;so that /C1030 /C12/C64x0 /C64x/C22/C64x0/C12 /C64x/C23dx/C22dx/C23ˆ/C103/C22/C23dx/C22dx/C23 or /C1030 /C12/C64x0 /C64x/C22/C64x0/C12 /C64x/C23ÿ/C103/C22/C23/C32! dx/C22dx/C23ˆ0: The above equation is identically zero for arbitrary dx/C22, so we have /C103/C22/C23ˆ/C64x0 /C64x/C22/C64x0/C12 /C64x/C23/C1030 /C12; …1:102† which shows that /C103/C22/C23is a covariant tensor of rank two. It is called the metric tensor or the fundamental tensor. Now contravariant and covariant tensors can be converted into each other with the help of the metric tensor. For example, we can get the covariant vector (tensor of rank one) A/C22from the contravariant vector A/C23: A/C22ˆ/C103/C22/C23A/C23: …1:103† Since we expect that the determinant of /C103/C22/C23does not vanish, the above equations can be solved for A/C23in terms of the A/C22. Let the result be A/C23ˆ/C103/C23/C22A/C22: …1:104† 52VECTOR AND TENSOR ANALYSIS By combining Eqs. (1.103) and (1.104) we get A/C22ˆ/C103/C22/C23/C103/C23 A : Since the equation must hold for any arbitrary A/C22,w eh a v e /C103/C22/C23/C103/C23 ˆ/C22 ; …1:105† where /C22 is Kronecker’s delta symbol. Thus, /C103/C22/C23is the inverse of /C103/C22/C23and vice versa; /C103/C22/C23is often called the conjugate or reciprocal tensor of /C103/C22/C23. But remember that /C103/C22/C23and/C103/C22/C23are the contravariant and covariant components of the same tensor, that is the metric tensor. Notice that the matrix ( /C103/C22/C23) is just the inverse of the matrix ( /C103/C22/C23). We can use /C103/C22/C23to lower any upper index occurring in a tensor, and use /C103/C22/C23to raise any lower index. It is necessary to remember the position from which the index was lowered or raised, because when we bring the index back to its original site, we do not want to interchange the order of indexes, in general T/C22/C236ˆT/C23/C22. Thus, for example A/C112 /C113ˆ/C103r/C112Ar/C113;A/C112/C113ˆ/C103r/C112/C103s/C113Ars;A/C112 rsˆ/C103r/C113A/C112/C113 s: /C65ssociated tensors All tensors obtained from a given tensor by forming an inner product with themetric tensor are called associated tensors of the given tensor. For example, A andA are associated tensors: A ˆ/C103 /C12A/C12;A ˆ/C103 /C12A/C12: /C71eodesics in a /C82iemannian space In a Euclidean space, the shortest path between two points is a straight linejoining the two points. In a Riemannian space, the shortest path between two points, called the geodesic, may be a curved path. To find the geodesic, let us consider a space curve in a Riemannian space given by x /C22ˆf/C22…t†and compute the distance between two points of the curve, which is given by the formula sˆZ/C81 P  /C103/C22dxdx/C22q ˆZt2 t1/C103 /C22d_xd_x/C22q dt; …1:106† where d_xˆdx=dt, and t(a parameter) varies from point to point of the geo- desic curve described by the relations which we are seeking. A geodesic joining 53ASSOCIATED TENSORS two points Pand/C81has a stationary value compared with any other neighboring path that connects Pand/C81. Thus, to find the geodesic we extremalize (1.106), and this leads to the di/C128erential equation of the geodesic (Problem 1.37) d dt/C64/C70 /C64_x ÿ/C64/C70 /C64xˆ0; …1:107† where /C70ˆ /C103 /C12_x _x/C12q ;and _xˆdx=dt. Now /C64/C70 /C64x/C13ˆ1 2/C103 /C12_x _x/C12ÿ1=2/C64/C103 /C12 /C64x/C13_x _x/C12;/C64/C70 /C64_x/C13ˆ12/C103 /C12_x _x/C12ÿ1=2 2/C103 /C13_x and ds=dtˆ /C103 /C12_x _x/C12q : Substituting these into (1.107) we obtain d dt/C103 /C13_x _sÿ1ÿ ÿ1 2/C64/C103 /C12 /C64x/C13_x _x/C12_sÿ1ˆ0;_sˆds dt or /C103 /C13/C127x ‡/C64/C103 /C13 /C64x/C12_x _x/C12ÿ12/C64/C103 /C12 /C64x/C13_x _x/C12ˆ/C103 /C13_x /C127s_sÿ1: We can simplify this equation by writing /C64/C103 /C13 /C64x/C12_x _x/C12ˆ1 2/C64/C103 /C13 /C64x/C12‡/C64/C103/C12/C13 /C64x  _x _x/C12; then we have /C103 /C13/C127x ‡ /C12; /C13‰Š _x _x/C12ˆ/C103 /C13_x /C127s_sÿ1: We can further simplify this equation by taking arc length as the parameter t, then _sˆ1;/C127sˆ0 and we have /C103 /C13d2x ds2‡ /C12; /C13‰Šdx dsdx/C12 dsˆ0: …1:108† where the functions ‰ /C12; /C13Šˆÿ /C12;/C13ˆ1 2/C64/C103 /C13 /C64x/C12‡/C64/C103/C12/C13 /C64x ÿ/C64/C103 /C12 /C64x/C13 …1:109† are called the Christo/C128el symbols of the first kind. Multiplying (1.108) by /C103/C26/C13, we obtain d2x/C26 ds2‡/C26 /C12() dx dsdx/C12 dsˆ0; …1:110† 54VECTOR AND TENSOR ANALYSIS where the functions /C26 /C12() ˆÿ/C26 /C12ˆ/C103/C26/C13‰ /C12; /C13Š… 1:111† are the Christo/C128el symbol of the second kind. Eq. (1.110) is, of course, a set of /C78coupled di/C128erential equations; they are the equations of the geodesic. In Euclidean spaces, geodesics are straight lines. In a Euclidean space, /C103 /C12are independent of the coordinates x/C22, so that the Christo/C128el symbols identically vanish, and Eq. (1.110) reduces to d2x/C26 ds2ˆ0 with the solution x/C26ˆa/C26s‡b/C26; where a/C26andb/C26are constants independent of s. This solution is clearly a straight line. The Christo/C128el symbols are not tensors. Using the defining Eqs. (1.109) and the transformation of the metric tensor, we can find the transformation laws of the Christo/C128el symbol. We now give the result, without the mathematical details: ÿ/C22/C23;ˆÿ /C12;/C13/C64x /C64x/C22/C64x/C12 /C64x/C23/C64x/C13 /C64x‡/C103 /C12/C64x /C64x/C642x/C12 /C64x/C22/C64x/C23: …1:112† The Christo/C128el symbols are not tensors because of the presence of the second term on the right hand side. /C67o/C118ariant di/C128erentiation We have seen that a covariant vector is transformed according to the formula A/C22ˆ/C64x/C23 /C64x/C22A/C23; where the coecients are functions of the coordinates, and so vectors at di/C128erent points transform di/C128erently. Because of this fact, dA/C22is not a vector, since it is the di/C128erence of vectors located at two (infinitesimally separated) points. We can verify this directly: /C64A/C22 /C64x/C13ˆ/C64A/C23 /C64x/C12/C64x/C23 /C64x/C22/C64x/C12 /C64x/C13‡A/C23/C642x/C23 /C64x/C22/C64x/C13; …1:113† 55COVARIANT DIFFERENTIATION which shows that /C64A=/C64x/C12are not the components of a tensor because of the second term on the right hand side. The same also applies to the di/C128erential of a contravariant vector. But we can construct a tensor by the following device. From Eq. (1.111) we have ÿ /C22/C13ˆÿ/C26 /C28/C64x /C64x/C22/C64x/C28 /C64x/C13/C64x /C64x/C26‡/C642x /C64x/C22/C64x/C13/C64x /C64x: …1:114† Multiplying (1.114) by A and subtracting from (1.113), we obtain /C64A/C22 /C64x/C13ÿA ÿ /C22/C13ˆ/C64A /C64x/C12ÿA/C26ÿ/C26 /C12/C64x /C64x/C22/C64x/C12 /C64x/C13: …1:115† If we define A ;/C12ˆ/C64A /C64x/C12ÿA/C26ÿ/C26 /C12; …1:116† then (1.115) can be rewritten as A/C22;/C13ˆA ;/C12/C64x /C64x/C22/C64x/C12 /C64x/C13; which shows that A ;/C12is a covariant tensor of rank 2. This tensor is called the covariant derivative of A with respect to x/C12. The semicolon denotes covariant di/C128erentiation. In a Cartesian coordinate system, the Christo/C128el symbols vanish, and so covariant di/C128erentiation reduces to ordinary di/C128erentiation. The contravariant derivative is found by raising the index which denotes di/C128er- entiation: A/C22;ˆ/C103 A/C22 ; : …1:117† We can similarly determine the covariant derivative of a tensor of arbitrary rank. In doing so we find the following simple rule helps greatly: To obtain the covariant derivative of the tensor T with respect to x/C22/C44 we add to the ordinary derivative /C64T =/C64x/C22for each covariant index /C23…T /C23:†a term ÿÿ /C22/C23T  :/C44 and for each contravariant index /C23…T/C23 †a term ‡ÿ /C23/C22T  .... Thus, T/C22/C23; ˆ/C64T/C22/C23 /C64x ÿÿ/C12 /C22 T/C12/C23ÿÿ/C12 /C23 T/C22/C12; T/C22 /C23; ˆ/C64T/C22 /C23 /C64x ÿÿ/C12 /C23 T/C22 /C12‡ÿ/C22 /C12 T/C12 /C23: The covariant derivatives of both the metric tensor and the Kronnecker delta are identically zero (Problem 1.38). 56VECTOR AND TENSOR ANALYSIS Problems 1.1. Given the vector /C65ˆ…2;2;ÿ1†and/C66ˆ…6;ÿ3;2†, determine: (a)6/C65ÿ3/C66,(b)A2‡B2,(c)/C65/C66,(d) the angle between /C65and/C66,(e) the direction cosines of /C65,(f) the component of /C66in the direction of /C65. 1.2. Find a unit vector perpendicular to the plane of /C65ˆ…2;ÿ6;ÿ3†and /C66ˆ…4;3;ÿ1†. 1.3. Prove that: (a) the median to the base of an isosceles triangle is perpendicular to the base; ( b) an angle inscribed in a semicircle is a right angle. 1.4. Given two vectors /C65ˆ…2;1;ÿ1†,/C66ˆ…1;ÿ1;2†find: ( a)/C65/C66, and ( b)a unit vector perpendicular to the plane containing vectors /C65and/C66. 1.5. Prove: ( a) the law of sines for plane triangles, and ( b) Eq. (1.16a). 1.6. Evaluate …2^e1ÿ3^e2†‰ …^e1‡^e2ÿ^e3†…3^e1ÿ^e3†Š. 1.7. ( a) Prove that a necessary and sucient condition for the vectors /C65,/C66and /C67to be coplanar is that /C65…/C66/C67†ˆ0: (b) Find an equation for the plane determined by the three points P1…2;ÿ1;1†,P2…3;2;ÿ1†andP3…ÿ1;3;2†. 1.8. ( a) Find the transformation matrix for a rotation of new coordinate system through an angle /C30about the x3…ˆz†-axis. (b) Express the vector /C65ˆ3^e1‡2^e2‡^e3in terms of the triad ^e0 1^e0 2^e0 3where thex0 1x0 2axes are rotated 45 8about the x3-axis (the x3-a n d x0 3-axes coinciding). 1.9. Consider the linear transformation A0 iˆP3 jˆ1^e0 i^ejAjˆP3jˆ1ijAj. Show, using the fact that the magnitude of the vector is the same in both systems, that X3 iˆ1ijikˆjk…j;kˆ1;2;3†: 1.10. A curve /C67is defined by the parametric equation r…u†ˆx1…u†^e1‡x2…u†^e2‡x3…u†^e3; where uis the arc length of C measured from a fixed point on /C67,a n d ris the position vector of any point on /C67; show that: (a)dr=duis a unit vector tangent to /C67; (b) the radius of curvature of the curve /C67is given by /C26ˆd2x1 du2/C32!2 ‡d2x2 du2/C32!2 ‡d2x3 du2/C32!22 435ÿ1=2 : 57PROBLEMS 1.11. ( a) Show that the acceleration aof a particle which travels along a space curve with velocity /C118is given by aˆd/C118 dt^T‡/C1182 /C26^/C78; where ^T,^/C78, and /C26are as defined in the text. (b) Consider a particle Pmoving on a circular path of radius rwith constant angular speed /C33ˆd=dt(Fig. 1.22). Show that the acceleration aof the particle is given by aˆÿ/C332r: 1.12. A particle moves along the curve x1ˆ2t2;x2ˆt2ÿ4t;x3ˆ3tÿ5, where t is the time. Find the components of the particle’s velocity and acceleration at time tˆ1 in the direction ^e1ÿ3^e2‡2^e3. 1.13. ( a) Find a unit vector normal to the surface x2 1‡x22ÿx3ˆ1 at the point P(1,1,1). (b) Find the directional derivative of /C30ˆx21x2x3‡4x1x23at (1, ÿ2;ÿ1) in the direction 2 ^e1ÿ^e2ÿ2^e3. 1.14. Consider the ellipse given by r1‡r2ˆconst :(Fig. 1.23). Show that r1andr2 make equal angles with the tangent to the ellipse. 1.15. Find the angle between the surfaces x2 1‡x22‡x23ˆ9a n d x3ˆx21‡x22ÿ3. at the point (2, ÿ1, 2). 1.16. ( a)I ffandgare di/C128erentiable scalar functions, show that /C114…f/C103†ˆf/C114/C103‡/C103/C114f: (b) Find /C114rifrˆ…x21‡x22‡x23†1=2. (c) Show that /C114rnˆnrnÿ2r. 1.17. Show that: (a)/C114…r=r3†ˆ0. Thus the divergence of an inverse-square force is zero. 58VECTOR AND TENSOR ANALYSIS Figure 1.22. Motion on a circle. (b)I ffis a di/C128erentiable function and /C65is a di/C128erentiable vector function, then /C114…f/C65†ˆ… /C114 f†/C65‡f…/C114 /C65†: 1.18. ( a) What is the divergence of a gradient/C63 (b) Show that /C1142…1=r†ˆ0. (c) Show that r… /C114 r†6 ˆ… r/C114†r. 1.19 Given /C114/C69ˆ0;/C114/C72ˆ0;/C114/C69ˆÿ/C64H=/C64t;/C114/C72ˆ/C64/C69=/C64t, show that /C69and/C72satisfy the wave equation /C1142uˆ/C642u=/C64t2. The given equations are related to the source-free Maxwell’s equations of electromagnetic theory, /C69and/C72are the electric field and magnetic field intensities. 1.20. ( a) Find constants a,b,csuch that /C65ˆ…x1‡2x2‡ax3†^e1‡…bx1ÿ3x2ÿx3†^e2‡…4x1‡cx2‡2x3†^e3 is irrotational. (b) Show that /C65can be expressed as the gradient of a scalar function. 1.21. Show that a cylindrical coordinate system is orthogonal.1.22. Find the volume element d/C86in: (a) cylindrical and (b) spherical coordinates. Hint: The volume element in orthogonal curvilinear coordinates is d/C86ˆ/C104 1/C1042/C1043du1du2du3ˆ/C64…x1;x2;x3† /C64…u1;u2;u3†/C12/C12/C12/C12/C12/C12/C12/C12du 1du2du3: 1.23. Evaluate the integralR…1;2† …0;1†…x2ÿy†dx‡…y2‡x†dyalong (a) a straight line from (0, 1) to (1, 2); (b) the parabola xˆt;yˆt2‡1; (c) straight lines from (0, 1) to (1, 1) and then from (1, 1) to (1, 2). 1.24. Evaluate the integralR…1;1† …0;0†…x2‡y2†dxalong (see Fig. 1.24): (a) the straight line yˆx, (b) the circle arc of radius 1 ( xÿ1†2‡y2ˆ1. 59PROBLEMS Figure 1.23. 1.25. Evaluate the surface integralR S/C65daˆR S/C65^nda, where /C65ˆx1x2^e1ÿ x2 1^e2‡…x1‡x2†^e3,Sis that portion of the plane 2 x1‡2x2‡x3ˆ6 included in the first octant. 1.26. Verify Gauss’ theorem for /C65ˆ…2x1ÿx3†^e1‡x21x2^e2ÿx1x23^e3taken over the region bounded by x1ˆ0;x1ˆ1;x2ˆ0;x2ˆ1;x3ˆ0;x3ˆ1. 1.28 Show that the electrostatic field intensity /C69…r†of a point charge /C81at the origin has an inverse-square dependence on r. 1.28. Show, by using Stokes’ theorem, that the gradient of a scalar field is irrota- tional: /C114… /C114 /C30…r†† ˆ 0: 1.29. Verify Stokes’ theorem for /C65ˆ…2x1ÿx2†^e1ÿx2x23^e2ÿx22x3^e3, where Sis the upper half surface of the sphere x2 1‡x22‡x23ˆ1 and ÿis its boundary (a circle in the x1x2plane of radius 1 with its center at the origin). 1.30. Find the area of the ellipse x1ˆacos;x2ˆbsin. 1.31. Show thatR Sr^ndaˆ0, where Sis a closed surface which encloses a volume /C86. 1.33. Starting with Eq. (1.95), express A0/C23in terms of A/C22. 1.33. Show that velocity and acceleration are contravariant vectors and that the gradient of a scalar field is a covariant vector. 1.34. The Cartesian components of the acceleration vector are axˆd2x dt2; ayˆd2y dt2; azˆd2z dt2: Find the component of the acceleration vector in the spherical polar co- ordinates. 1.35. Show that the property of symmetry (or anti-symmetry) with respect to indexes of a tensor is invariant under coordinate transformation. 1.36. A covariant tensor has components xy;2yÿz2;xzin rectangular coordi- nates, find its covariant components in spherical coordinates. 60VECTOR AND TENSOR ANALYSIS Figure 1.24. Paths for a path integral. 1.37. Prove that a necessary condition that IˆRt/C81 tP/C70…t;x;_x†dtbe an extremum (maximum or minimum) is that d dt/C64/C70 /C64_x ÿ/C64/C70 /C64xˆ0: 1.38. Show that the covariant derivatives of: ( a) the metric tensor, and ( b) the Kronecker delta are identically zero. 61PROBLEMS 2 Ordinary di/C128erential e/C113uations Physicists have a variety of reasons for studying di/C128erential equations: almost all the elementary and numerous of the advanced parts of theoretical physics are posed mathematically in terms of di/C128erential equations. We devote three chapters to di/C128erential equations. This chapter will be limited to ordinary di/C128erential equations that are reducible to a linear form. Partial di/C128erential equations and special functions of mathematical physics will be dealt with in Chapters 10 and 7. A di/C128erential equation is an equation that contains derivatives of an unknown function which expresses the relationship we seek. If there is only one independent variable and, as a consequence, total derivatives like dx=dt, the equation is called an ordinary di/C128erential equation (ODE). A partial di/C128erential equation (PDE) contains several independent variables and hence partial deriva- tives. Theorder of a di/C128erential equation is the order of the highest derivative appear- ing in the equation; its degree is the power of the derivative of highest order after the equation has been rationalized, that is, after fractional powers of all deriva-tives have been removed. Thus the equation d2y dx2‡3dy dx‡2yˆ0 is of second order and first degree, and d3y dx3ˆ  1‡…dy=dx†3q is of third order and second degree, since it contains the term ( d3y=dx3†2after it is rationalized. 62 A di/C128erential equation is said to be linear if each term in it is such that the dependent variable or its derivatives occur only once, and only to the first power. Thus d3y dx3‡ydy dxˆ0 is not linear, but x3d3y dx3‡exsinxdy dx‡yˆlnx is linear. If in a linear di/C128erential equation there are no terms independent of y, the dependent variable, the equation is also said to be homogeneous ; this would have been true for the last equation above if the ‘ln x’ term on the right hand side had been replaced by zero. A very important property of linear homogeneous equations is that, if we know two solutions y1andy2, we can construct others as linear combinations of them. This is known as the principle of superposition and will be proved later when wedeal with such equations. Sometimes di/C128erential equations look unfamiliar. A trivial change of variables can reduce a seemingly impossible equation into one whose type is readily recog-nizable. Many di/C128erential equations are very dicult to solve. There are only a rela- tively small number of types of di/C128erential equation that can be solved in closed form. We start with equations of first order. A first-order di/C128erential equation can always be solved, although the solution may not always be expressible in terms of familiar functions. A solution (or integral) of a di/C128erential equation is the relation between the variables, not involving di/C128erential coecients, which satisfies thedi/C128erential equation. The solution of a di/C128erential equation of order nin general involves narbitrary constants. First-order di/C128erential equations A di/C128erential equation of the general form dy dxˆÿf…x;y† /C103…x;y†;or/C103…x;y†dy‡f…x;y†dxˆ0 …2:1† is clearly a first-order di/C128erential equation. Separable variables Iff…x;y†and/C103…x;y†are reducible to P…x†and/C81…y†, respectively, then we have /C81…y†dy‡P…x†dxˆ0: …2:2† Its solution is found at once by integrating. 63FIRST-ORDER DIFFERENTIAL EQUATIONS The reader may notice that dy=dxhas been treated as if it were a ratio of dyand dx, that can be manipulated independently. Mathematicians may be unhappy about this treatment. But, if necessary, we can justify it by considering dyand dxto represent small finite changes yandx, before we have actually reached the limit where each becomes infinitesimal. Example 2.1 Consider the di/C128erential equation dy=dxˆÿy2ex: We can rewrite it in the following form ÿdy=y2ˆexdxwhich can be integrated separately giving the solution 1=yˆex‡c; where cis an integration constant. Sometimes when the variables are not separable a di/C128erential equation may be reduced to one in which they are separable by a change of variable. The general form of di/C128erential equation amenable to this approach is dy=dxˆf…ax‡by†; …2:3† where fis an arbitrary function and aandbare constants. If we let /C119ˆax‡by, then bdy=dxˆd/C119=dxÿa, and the di/C128erential equation becomes d/C119=dxÿaˆbf…/C119† from which we obtain d/C119 a‡bf…/C119†ˆdx in which the variables are separated. Example 2.2 Solve the equation dy=dxˆ8x‡4y‡…2x‡yÿ1†2: Solution: Letwˆ2x‡y, then dy=dxˆdw=dxÿ2, and the di/C128erential equation becomes d/C119=dx‡2ˆ4/C119‡…/C119ÿ1†2 or d/C119=‰4/C119‡…/C119ÿ1†2ÿ2Šˆdx: The variables are separated and the equation can be solved. 64ORDINARY DIFFERENTIAL EQUATIONS A homogeneous di/C128erential equation which has the general form dy=dxˆf…y=x†… 2:4† may also be reduced, by a change of variable, to one with separable variables. This can be illustrated by the following example: Example 2.3 Solve the equation dy dxˆy2‡xy x2: Solution: The right hand side can be rewritten as …y=x†2‡…y=x†, and hence is a function of the single variable /C118ˆy=x: We thus use /C118both for simplifying the right hand side of our equation, and also for rewriting dy=dxin terms of /C118and x. Now dy dxˆd dx…x/C118†ˆ/C118‡xd/C118 dx and our equation becomes /C118‡xd/C118 dxˆ/C1182‡/C118 from which we have d/C118 /C1182ˆdx x: Integration gives ÿ1 /C118ˆlnx‡corxˆAeÿx=y; where cand A…ˆeÿc†are constants. Sometimes a nearly homogeneous di/C128erential equation can be reduced to homogeneous form which can then be solved by variable separation. This canbe illustrated by the by the following: Example 2.4 Solve the equation dy=dxˆ…y‡xÿ5†=…yÿ3xÿ1†: 65FIRST-ORDER DIFFERENTIAL EQUATIONS Solution: Our equation would be homogeneous if it were not for the constants ÿ5 and ÿ1 in the numerator and denominator respectively. But we can eliminate them by a change of variable: x0ˆx‡ ;y0ˆy‡/C12; where and /C12are constants specially chosen in order to make our equation homogeneous: dy0=dx0ˆ…y0‡x0†=y0ÿ3x0: Note that dy0=dx0ˆdy=dx. Trivial algebra yields ˆÿ1;/C12ˆÿ4. Now let /C118ˆy0=x0, then dy0 dx0ˆd dx0…x0/C118†ˆ/C118‡x0d/C118 dx0 and our equation becomes /C118‡x0d/C118 dx0ˆ/C118‡1 /C118ÿ3;or/C118ÿ3 ÿ/C1182‡4/C118‡1d/C118ˆdx0 x0 in which the variables are separated and the equation can be solved by integra- tion. Example 2.5 Fall of a skydiver. Solution: Assuming the parachute opens at the beginning of the fall, there are two forces acting on the parachute: the downward force of gravity mg, and the upward force of air resistance kv2. If we choose a coordinate system that has yˆ0 at the earth’s surface and increases upward, then the equation of motion of the falling diver, according to Newton’s second law, is md/C118=dtˆÿm/C103‡k/C1182; where mis the mass, gthe gravitational acceleration, and ka positive constant. In general the air resistance is very complicated, but the power-law approximation is useful in many instances in which the velocity does not vary appreciably. Experiments show that for a subsonic velocity up to 300 m/s, the air resistance is approximately proportional to /C1182. The equation of motion is separable: md/C118 m/C103ÿk/C1182ˆdt 66ORDINARY DIFFERENTIAL EQUATIONS or, to make the integration easier d/C118 /C1182ÿ…m/C103=k†ˆÿk mdt: Now 1 /C1182ÿ…m/C103=k†ˆ1 …/C118‡/C118t†…/C118ÿ/C118t†ˆ1 2/C118t1 /C118ÿ/C118tÿ1 /C118‡/C118t ; where /C1182 tˆm/C103=k. Thus 1 2/C118td/C118 /C118ÿ/C118tÿd/C118 /C118ÿ/C118t ˆÿk mdt: Integrating yields 1 2/C118tln/C118ÿ/C118t /C118‡/C118t ˆÿk mt‡c; where cis an integration constant. Solving for /C118we finally obtain /C118…t†ˆ/C118t‰1‡Bexp…ÿ2/C103t=/C118t†Š 1ÿBexp…ÿ2/C103t=/C118t†; where Bˆexp…2/C118tC†. It is easy to see that as t!1 , exp( ÿ2/C103t=/C118t†!0, and so /C118!/C118t; that is, if he falls from a sucient height, the diver will eventually reach a constant velocity given by /C118t, the terminal velocity. To determine the constants of integration, we need to know the value of k, which is about 30 kg/m for the earth’s atmosphere and a standard parachute. Exact equations We may integrate Eq. (2.1) directly if its left hand side is the di/C128erential duof some function u…x;y†, in which case the solution is of the form u…x;y†ˆC …2:5† and Eq. (2.1) is said to be exact. A convenient test to see if Eq. (2.1) is exact isdoes /C64/C103…x;y† /C64xˆ/C64f…x;y† /C64y: …2:6† To see this, let us go back to Eq. (2.5) and we have d‰u…x;y†Š ˆ0: On performing the di/C128erentiation we obtain /C64u /C64xdx‡/C64u /C64ydyˆ0: …2:7† 67FIRST-ORDER DIFFERENTIAL EQUATIONS It is a general property of partial derivatives of any well-behaved function that the order of di/C128erentiation is immaterial. Thus we have /C64 /C64y/C64u /C64x ˆ/C64 /C64x/C64u /C64y : …2:8† Now if our di/C128erential equation (2.1) is of the form of Eq. (2.7), we must be able to identify f…x;y†ˆ/C64u=/C64xand /C103…x;y†ˆ/C64u=/C64y: …2:9† Then it follows from Eq. (2.8) that /C64/C103…x;y† /C64xˆ/C64f…x;y† /C64y; which is Eq. (2.6). Example 2.6 Show that the equation xdy=dx‡…x‡y†ˆ0 is exact and find its general solu- tion. Solution: We first write the equation in standard form …x‡y†dx‡xdyˆ0: Applying the test of Eq. (2.6) we notice that /C64f /C64yˆ/C64 /C64y…x‡y†ˆ1 and/C64/C103 /C64xˆ/C64x /C64xˆ1: Therefore the equation is exact, and the solution is of the form indicated by Eq. (2.7). From Eq. (2.9) we have /C64u=/C64xˆx‡y;/C64u=/C64yˆx; from which it follows that u…x;y†ˆx2=2‡xy‡/C104…y†;u…x;y†ˆxy‡k…x†; where /C104…y†andk…x†arise from integrating u…x;y†with respect to xandy, respec- tively. For consistency, we require that /C104…y†ˆ0a n d k…x†ˆx2=2: Thus the required solution is x2=2‡xyˆc: It is interesting to consider a di/C128erential equation of the type /C103…x;y†dy dx‡f…x;y†ˆk…x†; …2:10† 68ORDINARY DIFFERENTIAL EQUATIONS where the left hand side is an exact di/C128erential …d=dx†‰u…x;y†Š, and k…x†on the right hand side is a function of xonly. Then the solution of the di/C128erential equation can be written as u…x;y†ˆZ k…x†dx: …2:11† Alternatively Eq. (2.10) can be rewritten as /C103…x;y†dy dx‡‰f…x;y†ÿk…x†Š ˆ0: …2:10a† Since the left hand side of Eq. (2.10) is exact, we have /C64/C103=/C64xˆ/C64f=/C64y: Then Eq. (2.10a) is exact as well. To see why, let us apply the test for exactness for Eq. (2.10a) which requires /C64 /C64x‰/C103…x;y†Š ˆ/C64 /C64y‰f…x;y†ÿk…x†Š ˆ/C64 /C64y‰f…x;y†Š: Thus Eq. (2.10a) satisfies the necessary requirement for being exact. We can thus write its solution as U…x;y†ˆc; where /C64U /C64yˆ/C103…x;y†and/C64U /C64xˆf…x;y†ÿk…x†: Of course, the solution U…x;y†ˆcmust agree with Eq. (2.11). Integrating factors If a di/C128erential equation in the form of Eq. (2.1) is not already exact, it sometimescan be made so by multiplying by a suitable factor, called an integrating factor. Although an integrating factor always exists for each equation in the form of Eq. (2.1), it may be dicult to find it. However, if the equation is linear, that is, if can be written dy dx‡f…x†yˆ/C103…x†… 2:12† an integrating factor of the form expZ f…x†dx …2:13† is always available. It is easy to verify this. Suppose that R…x†is the integrating factor we are looking for. Multiplying Eq. (2.12) by /C82, we have Rdy dx‡Rf…x†yˆR/C103…x†;orRdy‡Rf…x†ydxˆR/C103…x†dx: 69FIRST-ORDER DIFFERENTIAL EQUATIONS The right hand side is already integrable; the condition that the left hand side of Eq. (2.12) be exact gives /C64 /C64y‰Rf…x†yŠˆ/C64R /C64x; which yields dR=dxˆRf…x†;ordR=Rˆf…x†dx; and integrating gives lnRˆZ f…x†dx from which we obtain the integrating factor /C82we were looking for RˆexpZ f…x†dx : It is now possible to write the general solution of Eq. (2.12). On applying the integrating factor, Eq. (2.12) becomes d…ye/C70† dxˆ/C103…x†e/C70; where /C70…x†ˆR f…x†dx. The solution is clearly given by yˆeÿ/C70Z e/C70/C103…x†dx‡C : Example 2.7Show that the equation xdy=dx‡2y‡x 2ˆ0 is not exact; then find a suitable integrating factor that makes the equation exact. What is the solution of thisequation/C63 Solution: We first write the equation in the standard form …2y‡x 2†dx‡xdyˆ0; then we notice that /C64 /C64y…2y‡x2†ˆ2a n d/C64 /C64xxˆ1; which indicates that our equation is not exact. To find the required integrating factor that makes our equation exact, we rewrite our equation in the form of Eq. (2.12): dy dx‡2y xˆÿx 70ORDINARY DIFFERENTIAL EQUATIONS from which we find f…x†ˆ1=x, and so the required integrating factor is expZ …1=x†dx ˆexp…lnx†ˆx: Applying this to our equation gives x2dy dx‡2xy‡x3ˆ0o rd dxx2y‡x4=4ÿ ˆ0 which integrates to x2y‡x4=4ˆc; or yˆcÿx4 4x2: Example 2.8 /C82Lcircuits: A typical /C82Lcircuit is shown in Fig. 2.1. Find the current I…t†in the circuit as a function of time t. Solution: We need first to establish the di/C128erential equation for the current flowing in the circuit. The resistance /C82and the inductance Lare both constant. The voltage drop across the resistance is I/C82, and the voltage drop across the inductance is LdI=dt. Kirchho/C128 ’s second law for circuits then gives LdI…t† dt‡RI…t†ˆ/C69…t†; which is in the form of Eq. (2.12), but with tas the independent variable instead of xand Ias the dependent variable instead of y. Thus we immediately have the general solution I…t†ˆ1 LeÿRt=LZ eRt=L/C69…t†dt‡keÿRt=L; 71FIRST-ORDER DIFFERENTIAL EQUATIONS Figure 2.1. /C82Lcircuit. where kis a constant of integration (in electric circuits, /C67is used for capacitance). Given Ethis equation can be solved for I…t†. If the voltage Eis constant, we obtain I…t†ˆ1 LeÿRt=L/C69L ReÿRt=L ‡keÿRt=Lˆ/C69 R‡keÿRt=L: Regardless of the value of k, we see that I…t†!/C69=Rast!1 : Setting tˆ0 in the solution, we find kˆI…0†ÿ/C69=R: Bernoulli’s equation Bernoulli’s equation is a non-linear first-order equation that occurs occasionally in physical problems: dy dx‡f…x†yˆ/C103…x†yn; …2:14† where nis not necessarily integer. This equation can be made linear by the substitution /C119ˆyawith suitably chosen. We find this can be achieved if ˆ1ÿn: /C119ˆy1ÿnoryˆ/C1191=…1ÿn†: This converts Bernoulli’s equation into d/C119 dx‡…1ÿn†f…x†/C119ˆ…1ÿn†/C103…x†; which can be made exact using the integrating factor exp …R …1ÿn†f…x†dx†. /C83econd-order equations /C119ith constant coe/C129cients The general form of the nth-order linear di/C128erential equation with constant coef- ficients is dny dxn‡/C1121dnÿ1y dxnÿ1‡‡ /C112nÿ1dy dx‡/C112nyˆ…Dn‡/C1121Dnÿ1‡‡ /C112nÿ1D‡/C112n†yˆf…x†; where /C1121;/C1122;...are constants, f…x†is some function of x, and Dd=dx.I f f…x†ˆ0, the equation is called homogeneous; otherwise it is called a non-homo- geneous equation. It is important to note that the symbol /C68is meaningless unless applied to a function of xand is therefore not a mathematical quantity in the usual sense. /C68is an operator. 72ORDINARY DIFFERENTIAL EQUATIONS Many of the di/C128erential equations of this type which arise in physical problems are of second order and we shall consider in detail the solution of the equation d2y dt2‡ady dt‡byˆ…D2‡aD‡b†yˆf…t†; …2:15† where aandbare constants, and tis the independent variable. As an example, the equation of motion for a mass on a spring is of the form Eq. (2.15), with a representing the friction, cbeing the constant of proportionality in Hooke’s law for the spring, and f…t†some time-dependent external force acting on the mass. Eq. (2.15) can also apply to an electric circuit consisting of an inductor, a resistor, a capacitor and a varying external voltage. The solution of Eq. (2.15) involves first finding the solution of the equation with f…t†replaced by zero, that is, d2y dt2‡ady dt‡byˆ…D2‡aD‡b†yˆ0; …2:16† this is called the reduced or homogeneous equation corresponding to Eq. (2.15). Nature of the solution of linear equations We now establish some results for linear equations in general. For simplicity, weconsider the second-order reduced equation (2.16). If y 1andy2are independent solutions of (2.16) and Aand Bare any constants, then D…Ay1‡By2†ˆADy 1‡BDy 2;D2…Ay1‡By2†ˆAD2y1‡BD2y2 and hence …D2‡aD‡b†…Ay1‡By2†ˆA…D2‡aD‡b†y1‡B…D2‡aD‡b†y2ˆ0: Thus yˆAy1‡By2is a solution of Eq. (2.16), and since it contains two arbitrary constants, it is the general solution. A necessary and sucient condition for two solutions y1andy2to be linearly independent is that the Wronskian determinant of these functions does not vanish: y1y2 dy1 dtdy2 dt/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C126ˆ0: Similarly, if y 1;y2;...;ynarenlinearly independent solutions of the nth-order linear equations, then the general solution is yˆA1y1‡A2y2‡‡ Anyn; where A1;A2;...;Anare arbitrary constants. This is known as the superposition principle. 73SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS General solutions of the second-order equations Suppose that we can find one solution, y/C112…t†say, of Eq. (2.15): …D2‡aD‡b†y/C112…t†ˆf…t†: …2:15a† Then on defining yc…t†ˆy…t†ÿy/C112…t† we find by subtracting Eq. (2.15a) from Eq. (2.15) that …D2‡aD‡b†yc…t†ˆ0: That is, yc…t†satisfies the corresponding homogeneous equation (2.16), and it is known as the complementary function yc…t†of non-homogeneous equation (2.15). while the solution y/C112…t†is called a particular integral of Eq. (2.15). Thus, the general solution of Eq. (2.15) is given by y…t†ˆAyc…t†‡By/C112…t†: …2:17† Finding the complementary function Clearly the complementary function is independent of f…t†, and hence has nothing to do with the behavior of the system in response to the external applied influence. What it does represent is the free motion of the system. Thus, for example, even without external forces applied, a spring can oscillate, because of any initial displacement and/or velocity. Similarly, had a capacitor already been charged attˆ0, the circuit would subsequently display current oscillations even if there is no applied voltage. In order to solve Eq. (2.16) for yc…t†, we first consider the linear first-order equation ady dt‡byˆ0: Separating the variables and integrating, we obtain yˆAeÿbt=a; where Ais an arbitrary constant of integration. This solution suggests that Eq. (2.16) might be satisfied by an expression of the type yˆe/C112t; where pis a constant. Putting this into Eq. (2.16), we have e/C112t…/C1122‡a/C112‡b†ˆ0: Therefore yˆe/C112tis a solution of Eq. (2.16) if /C1122‡a/C112‡bˆ0: 74ORDINARY DIFFERENTIAL EQUATIONS This is called the auxiliary (or characteristic) equation of Eq. (2.16). Solving it gives /C1121ˆÿa‡ a2ÿ4bp 2; /C1122ˆÿaÿa2ÿ4bp 2: …2:18† We now distinguish between the cases in which the roots are real and distinct, complex or coincident. (i) Real and distinct roots ( a2ÿ4b/C620† In this case, we have two independent solutions y1ˆe/C1121t;y2ˆe/C1122tand the general solution of Eq. (2.16) is a linear combination of these two: yˆAe/C1121t‡Be/C1122t; …2:19† where Aand Bare constants. Example 2.9 Solve the equation …D2ÿ2Dÿ3†yˆ0, given that yˆ1 and y0ˆdy=dxˆ2 when tˆ0. Solution: The auxiliary equation is p2ÿ2pÿ3ˆ0, from which we find pˆÿ1 orpˆ3. Hence the general solution is yˆAeÿt‡Be3t: The constants Aand Bcan be determined by the boundary conditions at tˆ0. Since yˆ1 when tˆ0, we have 1ˆA‡B: Now y0ˆÿAeÿt‡3Be3t and since y0ˆ2 when tˆ0, we have 2 ˆÿA‡3B. Hence Aˆ1=4;Bˆ3=4 and the solution is 4yˆeÿt‡3e3t: (ii) Complex roots …a2ÿ4b<0† If the roots /C1121,/C1122of the auxiliary equation are imaginary, the solution given by Eq. (2.18) is still correct. In order to give the solutions in terms of real quantities,we can use the Euler relations to express the exponentials. If we let rˆÿa=2;isˆ a 2ÿ4bp =2, then e/C1121tˆerteistˆert‰cosst‡isinstŠ; e/C1122tˆerteistˆert‰cosstÿisinstŠ 75SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS and the general solution can be written as yˆAe/C1121t‡Be/C1122t ˆert‰…A‡B†cosst‡i…AÿB†sinstŠ ˆert‰A0cosst‡B0sinstŠ… 2:20† with A0ˆA‡B;B0ˆi…AÿB†: The solution (2.20) may be expressed in a slightly di/C128erent and often more useful form by writing B0=A0ˆtan. Then yˆ…A2 0‡B20†1=2ert…coscosst‡sinsinst†ˆCertcos…stÿ†; …2:20a† where /C67andare arbitrary constants. Example 2.10 Solve the equation …D2‡4D‡13†yˆ0, given that yˆ1 and y0ˆ2 when tˆ0. Solution: The auxiliary equation is p2‡4p‡13ˆ0, and hence pˆÿ23i. The general solution is therefore, from Eq. (2.20), yˆeÿ2t…A0cos 3 t‡B0sin 3t†: Since yˆ/C108when tˆ0, we have A0ˆ1. Now y0ˆÿ2eÿ2t…A0cos 3 t‡B0sin 3t†‡3eÿ2t…ÿA0sin 3t‡B0cos 3 t† and since y0ˆ2 when tˆ0, we have 2 ˆÿ2A0‡3B0. Hence B0ˆ4=3, and the solution is 3yˆeÿ2t…3 cos 3 t‡4 sin 3 t†: (iii) Coincident roots When a2ˆ4b, the auxiliary equation yields only one value for p, namely /C112ˆ ˆÿa=2, and hence the solution yˆAe t. This is not the general solution as it does not contain the necessary two arbitrary constants. In order to obtain thegeneral solution we proceed as follows. Assume that yˆ/C118e t, where vis a func- tion of tto be determined. Then y0ˆ/C1180e t‡ /C118e t;y00ˆ/C11800e t‡2 /C1180e t‡ 2/C118e t: Substituting for y;y0, and y00in the di/C128erential equation we have e t‰/C11800‡2 /C1180‡ 2/C118‡a…/C1180‡ /C118†‡b/C118Šˆ0 and hence /C11800‡/C1180…a‡2 †‡/C118… 2‡a ‡b†ˆ0: 76ORDINARY DIFFERENTIAL EQUATIONS Now 2‡a ‡bˆ0;and a‡2 ˆ0 so that /C11800ˆ0: Hence, integrating gives /C118ˆAt‡B; where Aand Bare arbitrary constants, and the general solution of Eq. (2.16) is yˆ…At‡B†e t…2:21† Example 2.11 Solve the equation ( D2ÿ4D‡4†yˆ0 given that yˆ1a n d Dyˆ3 when tˆ0: Solution: The auxiliary equation is p2ÿ4p‡4ˆ…pÿ2†2ˆ0 which has one root pˆ2. The general solution is therefore, from Eq. (2.21) yˆ…At‡B†e2t: Since yˆ1 when tˆ0, we have Bˆ1. Now y0ˆ2…At‡B†e2t‡Ae2t and since Dyˆ3 when tˆ0, 3ˆ2B‡A: Hence Aˆ1 and the solution is yˆ…t‡1†e2t: Finding the particular integral The particular integral is a solution of Eq. (2.15) that takes the term f…t†on the right hand side into account. The complementary function is transient in nature,so from a physical point of view, the particular integral will usually dominate the response of the system at large times. The method of determining the particular integral is to guess a suitable func- tional form containing arbitrary constants, and then to choose the constants to ensure it is indeed the solution. If our guess is incorrect, then no values of these constants will satisfy the di/C128erential equation, and so we have to try a di/C128erent form. Clearly this procedure could take a long time; fortunately, there are some guiding rules on what to try for the common examples of f(t): 77SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS (1)f…t†ˆa polynomial in t. Iff…t†is a polynomial in twith highest power tn, then the trial particular integral is also a polynomial in t, with terms up to the same power. Note that the trial particular integral is a power series in t, even if f…t†contains only a single terms Atn. (2)f…t†ˆAekt. The trial particular integral is yˆBekt. (3)f…t†ˆAsinktorAcoskt. The trial particular integral is yˆAsinkt‡Ccoskt. That is, even though f…t†contains only a sine or cosine term, we need both sine and cosine terms for the particular integral. (4)f…t†ˆAe tsin/C12torAe tcos/C12t. The trial particular integral is yˆe t…Bsin/C12t‡Ccos/C12t†. (5)f…t†is a polynomial of order nint, multiplied by ekt. The trial particular integral is a polynomial in twith coecients to be determined, multiplied by ekt. (6)f…t†is a polynomial of order nint, multiplied by sin kt. The trial particular integral is yˆ/C6n jˆ0…Bjsinkt‡Cjcoskt†tj. Can we try yˆ…Bsinkt‡Ccoskt†/C6njˆ0Djtj/C63 The answer is no. Do you know why/C63 If the trial particular integral or part of it is identical to one of the terms of the complementary function, then the trial particular integral must be multiplied by an extra power of t. Therefore, we need to find the complementary function before we try to work out the particular integral. What do we mean by ‘identical inform’/C63 It means that the ratio of their t-dependences is a constant. Thus ÿ2e ÿt andAeÿtare identical in form, but eÿtandeÿ2tare not. Particular integral and the operator /C68…ˆd=dx† We now describe an alternative method that can be used for finding particular integrals. As compared with the method described in previous section, it involves less guesswork as to what the form of the solution is, and the constants multi-plying the functional forms of the answer are obtained automatically. It does, however, require a fair amount of practice to ensure that you are familiar with how to use it. The technique involves using the di/C128erential operator Dd…†=dt, which is an interesting and simple example of a linear operator without a matrix representa- tion. It is obvious that /C68obeys the relevant laws of operator algebra: suppose f and gare functions of t, and ais a constant, then (i)D…f‡/C103†ˆDf‡D/C103 (distributive); (ii)DafˆaDf (commutative); (iii)D nDmfˆDn‡mf (index law). 78ORDINARY DIFFERENTIAL EQUATIONS We can form a polynomial function of /C68and write /C70…D†ˆa0Dn‡a1Dnÿ1‡‡ anÿ1D‡an so that /C70…D†f…t†ˆa0Dnf‡a1Dnÿ1f‡‡ anÿ1Df‡anf and we can interpret Dÿ1as follows Dÿ1Df…t†ˆf…t† andZ …Df†dtˆf: Hence Dÿ1indicates the operation of integration (the inverse of di/C128erentiation). Similarly Dÿmfmeans ‘integrate f…t†mtimes’. These properties of the linear operator /C68can be used to find the particular integral of Eq. (2.15): d2y dt2‡ady dt‡byˆD2‡aD‡bÿ yˆf…t† from which we obtain yˆ1 D2‡aD‡bf…t†ˆ1 /C70…D†f…t†; …2:22† where /C70…D†ˆD2‡aD‡b: The trouble with Eq. (2.22) is that it contains an expression involving /C68s in the denominator. It requires a fair amount of practice to use Eq. (2.22) to express yin terms of conventional functions. For this, there are several rules to help us. Rules for /C68operators Given a power series of /C68 /C71…D†ˆa0‡a1D‡‡ anDn‡ and since Dne tˆ ne t, it follows that /C71…D†e tˆ…a0‡a1D‡‡ anDn‡ † e tˆ/C71… †e t: Thus we have Rule (a): /C71…D†e tˆ/C71… †e tprovided /C71… †is convergent. When /C71…D†is the expansion of 1 =/C70…D†this rule gives 1 /C70…D†e tˆ1 /C70… †e tprovided /C70… †6 ˆ0: 79SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS Now let us operate /C71…D†on a product function e t/C86…t†: /C71…D†‰e t/C86…t†Š ˆ ‰ /C71…D†e tŠ/C86…t†‡e t‰/C71…D†/C86…t†Š ˆe t‰/C71… †‡/C71…D†Š/C86…t†ˆe t/C71…D‡ †‰/C86…t†Š: That is, we have Rule (b): /C71…D†‰e t/C86…t†Š ˆe t/C71…D‡ †‰/C86…t†Š: Thus, for example D2‰e tt2Šˆe t…D‡ †2‰t2Š: Rule (c): /C71…D2†sinktˆ/C71…ÿk2†sinkt: Thus, for example 1 D2…sin 3t†ˆÿ1 9sin 3t: Example 2.12 /C68amped oscillations (Fig. 2.2) Suppose we have a spring of natural length L(that is, in its unstretched state). If we hang a ball of mass mfrom it and leave the system in equilibrium, the spring stretches an amount d, so that the ball is now L‡dfrom the suspension point. We measure the vertical displacement of the ball from this static equilibrium point. Thus, L‡disyˆ0, and yis chosen to be positive in the downward direction, and negative upward. If we pull down on the ball and then release it, it oscillates up and down about the equilibrium position. To analyze the oscilla- tion of the ball, we need to know the forces acting on it: 80ORDINARY DIFFERENTIAL EQUATIONS Figure 2.2. Damped spring system. (1) the downward force of gravity, mg: (2) the restoring force kywhich always opposes the motion (Hooke’s law), where kis the spring constant of the spring. If the ball is pulled down a distance yfrom its static equilibrium position, this force is ÿk…d‡y†. Thus, the total net force acting on the ball is m/C103ÿk…d‡y†ˆm/C103ÿkdÿky: In static equilibrium, yˆ0 and all forces balances. Hence kdˆm/C103 and the net force acting on the spring is just ÿky; and the equation of motion of the ball is given by Newton’s second law of motion: md2y dt2ˆÿky; which describes free oscillation of the ball. If the ball is connected to a dashpot (Fig. 2.2), a damping force will come into play. Experiment shows that the damp- ing force is given by ÿbdy=dt, where the constant bis called the damping constant. The equation of motion of the ball now is md2y dt2ˆÿkyÿbdy dtory00‡b my0‡k myˆ0: The auxiliary equation is /C1122‡b m/C112‡k mˆ0 with roots /C1121ˆÿb 2m‡1 2m b2ÿ4km/C112 ; /C1122ˆÿb 2mÿ1 2mb 2ÿ4km/C112 : We now have three cases, resulting in quite di/C128erent motions of the oscillator. /C67ase 1 b2ÿ4km/C620 (overdamping) The solution is of the form y…t†ˆc1e/C1121t‡c2e/C1122t: Now, both band kare positive, so 1 2m b2ÿ4km/C112 <b 2m and accordingly /C1121ˆÿb 2m‡1 2m b2ÿ4km/C112 <0: 81SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS Obviously /C1122<0 also. Thus, y…t†!0a s t!1 . This means that the oscillation dies out with time and eventually the mass will assume the static equilibrium position. /C67ase 2 b2ÿ4kmˆ0 (critical damping) The solution is of the form y…t†ˆeÿbt=2m…c1‡c2t†: As both bandmare positive, y…t†!0a st!1 as in case 1. But c1andc2play a significant role here. Since eÿbt=2m6ˆ0 for finite t,y…t†can be zero only when c1‡c2tˆ0, and this happens when tˆÿc1=c2: If the number on the right is positive, the mass passes through the equilibrium position yˆ0 at that time. If the number on the right is negative, the mass never passes through the equilibrium position. It is interesting to note that c1ˆy…0†, that is, c1measures the initial position. Next, we note that y0…0†ˆc2ÿbc1=2m;orc2ˆy0…0†‡by…0†=2m: /C67ase 3 b2ÿ4km<0 (underdamping) The auxiliary equation now has complex roots /C1121ˆÿb 2m‡i 2m 4kmÿb2/C112 ;/C1122ˆÿb 2mÿi 2m4kmÿb 2/C112 and the solution is of the form y…t†ˆeÿbt=2mc1cos4kmÿb 2/C112t 2m‡c2sin4kmÿb 2/C112t 2m /C104/C105 ; which can be rewritten as y…t†ˆceÿbt=2mcos…/C33tÿ †; where cˆ c2 1‡c22q ; ˆtanÿ1c2 c1 ;and /C33ˆ 4kmÿb2/C112 =2m: As in case 2, eÿbt=2m!0a s t!1 , and the oscillation gradually dies down to zero with increasing time. As the oscillator dies down, it oscillates with a fre- quency /C33=2. But the oscillation is not periodic. 82ORDINARY DIFFERENTIAL EQUATIONS /C84he /C69uler linear equation The linear equation with variable coecients xndny dxn‡/C1121xnÿ1dnÿ1y dxnÿ1‡‡ /C112nÿ1xdy dx‡/C112nyˆf…x†; …2:23† in which the derivative of the jth order is multiplied by xjand by a constant, is known as the Euler or Cauchy equation. It can be reduced, by the substitution xˆet, to a linear equation with constant coecients with tas the independent variable. Now if xˆet, then dx=dtˆx,a n d dy dxˆdy dtdt dxˆ1 xdy dt;orxdy dxˆdy dt and d2y dx2ˆd dxdy dx ˆd dt1 xdy dtdt dxˆ1 xd dt1 xdy dt or xd2y dx2ˆ1 xd2y dt2‡dy dtd dt1 x ˆ1 xd2y dx2ÿ1 xdy dt and hence x2d2y dx2ˆd2y dt2ÿdy dtˆd dtdy dtÿ1 y: Similarly x3d3y dx3ˆd dtd dtÿ1d dtÿ2 y; and xndny dxnˆd dtd dtÿ1d dtÿ2 d dtÿn‡1 y: Substituting for xj…djy=dxj†in Eq. (2.23) the equation transforms into dny dtn‡/C1131dnÿ1y dtnÿ1‡‡ /C113nÿ1dy dt‡/C113nyˆf…et† in which /C1131,/C1132;...;/C113nare constants. Example 2.13 Solve the equation x2d2y dx2‡6xdy dx‡6yˆ1 x2: 83THE EULER LINEAR EQUATION Solution: Putxˆet, then xdy dxˆdy dt;x2d2y dx2ˆd2y dt2ÿdy dt: Substituting these in the equation gives d2y dt2‡5dy dt‡6yˆet: The auxiliary equation /C1122‡5/C112‡6ˆ…/C112‡2†…/C112‡3†ˆ0 has two roots: /C1121ˆÿ2, /C1122ˆ3. So the complementary function is of the form ycˆAeÿ2t‡Beÿ3tand the particular integral is y/C112ˆ1 …D‡2†…D‡3†eÿ2tˆteÿ2t: The general solution is yˆAeÿ2t‡Beÿ3t‡teÿ2t: The Euler equation is a special case of the general linear second-order equation D2y‡/C112…x†Dy‡/C113…x†yˆf…x†; where /C112…x†,/C113…x†, and f…x†are given functions of x. In general this type of equation can be solved by series approximation methods which will be introduced in next section, but in some instances we may solve it by means of a variable substitution, as shown by the following example: D2y‡…4xÿxÿ1†Dy‡4x2yˆ0; where /C112…x†ˆ… 4xÿxÿ1†;/C113…x†ˆ4x2;and f…x†ˆ0: If we let xˆz1=2 the above equation is transformed into the following equation with constant coecients: D2y‡2Dy‡yˆ0; which has the solution yˆ…A‡Bz†eÿz: Thus the general solution of the original equation is yˆ…A‡Bx2†eÿx2: 84ORDINARY DIFFERENTIAL EQUATIONS /C83olutions in po/C119er series In many problems in physics and engineering, the di/C128erential equations are of such a form that it is not possible to express the solution in terms of elementary functions such as exponential, sine, cosine, etc.; but solutions can be obtained as convergent infinite series. What is the basis of this method/C63 To see it, let us consider the following simple second-order linear di/C128erential equation d2y dx2‡yˆ0: Now assuming the solution is given by yˆa0‡a1x‡a2x2‡ , we further assume the series is convergent and di/C128erentiable term by term for suciently small x. Then dy=dxˆa1‡2a2x‡3a3x2‡ and d2y=dx2ˆ2a2‡23a3x‡34a4x2‡ : Substituting the series for yandd2y=dx2in the given di/C128erential equation and collecting like powers of xyields the identity …2a2‡a0†‡…23a3‡a1†x‡…34a4‡a2†x2‡ˆ 0: Since if a power series is identically zero all of its coecients are zero, equating to zero the term independent of xand coecients of x,x2;...;gives 2a2‡a0ˆ0; 45a5‡a3ˆ0; 23a3‡a1ˆ0; 56a6‡a4ˆ0; 34a4‡a2ˆ0;  and it follows that a2ˆÿa0 2;a3ˆÿa1 23ˆÿa1 3/C33;a4ˆÿa2 34ÿa0 4/C33 a5ˆÿa3 45ˆa1 5/C33;a6ˆÿa4 56ˆÿa0 6/C33;...: The required solution is yˆa01ÿx2 2/C33‡x4 4/C33ÿx6 6/C33‡ÿ/C32! ‡a1xÿx3 3/C33‡x5 5/C33ÿ‡/C32! ; you should recognize this as equivalent to the usual solutionyˆa 0cosx‡a1sinx,a0anda1being arbitrary constants. 85SOLUTIONS IN POWER SERIES Ordinary and singular points of a di/C128erential equation We shall concentrate on the linear second-order di/C128erential equation of the form d2y dx2‡P…x†dy dx‡/C81…x†yˆ0 …2:24† which plays a very important part in physical problems, and introduce certain definitions and state (without proofs) some important results applicable to equa- tions of this type. With some small modifications, these are applicable to linear equation of any order. If both the functions Pand/C81can be expanded in Taylor series in the neighborhood of xˆ , then Eq. (2.24) is said to possess an ordinary point at xˆ . But when either of the functions Por/C81does not possess a Taylor series in the neighborhood of xˆ , Eq. (2.24) is said to have a singular point at xˆ .I f Pˆ…x†=…xÿ †and /C81ˆ/C22…x†=…xÿ †2 and…x†and/C22…x†can be expanded in Taylor series near xˆ . In such cases, xˆ is a singular point but the singularity is said to be regular. Frobenius and Fuchs theorem Frobenius and Fuchs showed that: (1) If P…x†and/C81…x†are regular at xˆ , then the di/C128erential equation (2.24) possesses two distinct solutions of the form yˆX1 ˆ0a…xÿ †…a06ˆ0†: …2:25† (2) If P…x†and/C81…x†are singular at xˆ , but …xÿ †P…x†and…xÿ †2/C81…x† are regular at xˆ , then there is at least one solution of the di/C128erential equation (2.24) of the form yˆX1 ˆ0a…xÿ †‡/C26…a06ˆ0†; …2:26† where /C26is some constant, which is valid for jxÿ j</C12whenever the Taylor series for …x†and/C22…x†are valid for these values of x. (3) If P…x†and/C81…x†are irregular singular at xˆ (that is, …x†and/C22…x†are singular at xˆ †, then regular solutions of the di/C128erential equation (2.24) may not exist. 86ORDINARY DIFFERENTIAL EQUATIONS The proofs of these results are beyond the scope of the book, but they can be found, for example, in E. L. Ince’s Ordinary /C68i/C128erential E/C113uations , Dover Publications Inc., New York, 1944. The first step in finding a solution of a second-order di/C128erential equation relative to a regular singular point xˆ is to determine possible values for the index /C26in the solution (2.26). This is done by substituting series (2.26) and its appropriate di/C128erential coecients into the di/C128erential equation and equating to zero the resulting coecient of the lowest power of xÿ .This leads to a quadratic equation, called the indicial equation, from which suitable values of /C26can be found. In the simplest case, these values of /C26will give two di/C128erent series solutions and the general solution of the di/C128erential equation is then given by a linear combination of the separate solutions. The complete procedure is shown in Example 2.14 below. Example 2.14 Find the general solution of the equation 4xd2y dx2‡2dy dx‡yˆ0: Solution: The origin is a regular singular point and, writing yˆ/C61 ˆ0ax‡/C26…a06ˆ0†we have dy=dxˆX1 ˆ0a…‡/C26†x‡/C26ÿ1;d2y=dx2ˆX1 ˆ0a…‡/C26†…‡/C26ÿ1†x‡/C26ÿ2: Before substituting in the di/C128erential equation, it is convenient to rewrite it in the form 4xd2y dx2‡2dy dx() ‡yfgˆ0: When ax‡/C26is substituted for y, each term in the first bracket yields a multiple of x‡/C26ÿ1, while the second bracket gives a multiple of x‡/C26and, in this form, the di/C128erential equation is said to be arranged according to weight, the weights of the bracketed terms di/C128ering by unity. When the assumed series and its di/C128erential coecients are substituted in the di/C128erential equation, the term containing the lowest power of xis obtained by writing yˆa0x/C26in the first bracket. Since the coecient of the lowest power of xmust be zero and, since a06ˆ0, this gives the indicial equation 4/C26…/C26ÿ1†‡2/C26ˆ2/C26…2/C26ÿ1†ˆ0; its roots are /C26ˆ0,/C26ˆ1=2. 87SOLUTIONS IN POWER SERIES The term in x‡/C26is obtained by writing yˆa‡1x‡/C26‡1in first bracket and yˆax‡/C26in the second. Equating to zero the coecient of the term obtained in this way we have f4…‡/C26‡1†…‡/C26†‡2…‡/C26‡1†ga‡1‡aˆ0; giving, with replaced by n, an‡1ˆÿ1 2…/C26‡n‡1†…2/C26‡2n‡1†an: This relation is true for nˆ1;2;3;...and is called the recurrence relation for the coecients. Using the first root /C26ˆ0 of the indicial equation, the recurrence relation gives an‡1ˆ1 2…n‡1†…2n‡1†an and hence a1ˆÿa0 2;a2ˆÿa1 12ˆa0 4/C33;a3ˆÿa2 30ˆÿa0 6/C33;...: Thus one solution of the di/C128erential equation is the series a01ÿx 2/C33‡x2 4/C33ÿx3 6/C33‡ÿ/C32! : With the second root /C26ˆ1=2, the recurrence relation becomes an‡1ˆÿ1 …2n‡3†…2n‡2†an: Replacing a0(which is arbitrary) by b0, this gives a1ˆÿb0 32ˆÿb0 3/C33;a2ˆÿa1 54ˆb0 5/C33;a3ˆÿa2 76ˆÿb0 7/C33;...: and a second solution is b0x1=21ÿx 3/C33‡x2 5/C33ÿx3 7/C33‡ÿ/C32! : The general solution of the equation is a linear combination of these two solu- tions. Many physical problems require solutions which are valid for large values of the independent variable x. By using the transformation xˆ1=t, the di/C128erential equation can be transformed into a linear equation in the new variable tand the solutions required will be those valid for small t. In Example 2.14 the indicial equation has two distinct roots. But there are two other possibilities: ( a) the indicial equation has a double root; ( b) the roots of the 88ORDINARY DIFFERENTIAL EQUATIONS indicial equation di/C128er by an integer. We now take a general look at these cases. For this purpose, let us consider the following di/C128erential equation which is highlyimportant in mathematical physics: x 2y00‡x/C103…x†y0‡/C104…x†yˆ0; …2:27† where the functions /C103…x†and/C104…x†are analytic at xˆ0. Since the coecients are not analyic at xˆ0, the solution is of the form y…x†ˆxrX1 mˆ0amxm…a06ˆ0†: …2:28† We first expand /C103…x†and/C104…x†in power series, /C103…x†ˆ/C1030‡/C1031x‡/C1032x2‡ /C104…x†ˆ/C1040‡/C1041x‡/C1042x2‡ : Then di/C128erentiating Eq. (2.28) term by term, we find y0…x†ˆX1 mˆ0…m‡r†amxm‡rÿ1;y00…x†ˆX1 mˆ0…m‡r†…m‡rÿ1†amxm‡rÿ2: By inserting all these into Eq. (2.27) we obtain xr‰r…rÿ1†a0‡ Ї… /C1030‡/C1031x‡ † xr…ra0‡ † ‡…/C1040‡/C1041x‡ † xr…a0‡a1x‡ †ˆ 0: Equating the sum of the coecients of each power of xto zero, as before, yields a system of equations involving the unknown coecients am. The smallest power is xr, and the corresponding equation is ‰r…rÿ1†‡/C1030r‡/C1040Ša0ˆ0: Since by assumption a06ˆ0, we obtain r…rÿ1†‡/C1030r‡/C1040ˆ0o r r2‡…/C1030ÿ1†r‡/C1040ˆ0: …2:29† This is the indicial equation of the di/C128erential equation (2.27). We shall see thatour series method will yield a fundamental system of solutions; one of the solu- tions will always be of the form (2.28), but for the form of other solution there will be three di/C128erent possibilities corresponding to the following cases. Case 1 The roots of the indicial equation are distinct and do not di/C128er by an integer. Case 2 The indicial equation has a double root. Case 3 The roots of the indicial equation di/C128er by an integer. We now discuss these cases separately. 89SOLUTIONS IN POWER SERIES /C67ase 1 /C68istinct roots not di/C128ering by an integer This is the simplest case. Let r1andr2be the roots of the indicial equation (2.29). If we insert rˆr1into the recurrence relation and determine the coecients a1,a2;...successively, as before, then we obtain a solution y1…x†ˆxr1…a0‡a1x‡a2x2‡ † : Similarly, by inserting the second root rˆr2into the recurrence relation, we will obtain a second solution y2…x†ˆxr2…a0/C42‡a1/C42x‡a2/C42x2‡ † : Linear independence of y1andy2follows from the fact that y1=y2is not constant because r1ÿr2is not an integer. /C67ase 2 /C68ouble rootsThe indicial equation (2.29) has a double root rif, and only if, …/C103 0ÿ1†2ÿ4/C1040ˆ0, and then rˆ…1ÿ/C1030†=2. We may determine a first solution y1…x†ˆxr…a0‡a1x‡a2x2‡ † rˆ1ÿ/C1030 2 …2:30† as before. To find another solution we may apply the method of variation of parameters, that is, we replace constant cin the solution cy1…x†by a function u…x†to be determined, such that y2…x†ˆu…x†y1…x†… 2:31† is a solution of Eq. (2.27). Inserting y2and the derivatives y0 2ˆu0y1‡uy0 1 y00 2ˆu00y1‡2u0y0 1‡uy00 1 into the di/C128erential equation (2.27) we obtain x2…u00y1‡2u0y0 1‡uy00 1†‡x/C103…u0y1‡uy0 1†‡/C104uy 1ˆ0 or x2y1u00‡2x2y0 1u0‡x/C103y 1u0‡…x2y00 1‡x/C103y0 1‡/C104y1†uˆ0: Since y1is a solution of Eq. (2.27), the quantity inside the bracket vanishes; and the last equation reduces to x2y1u00‡2x2y0 1u0‡x/C103y 1u0ˆ0: Dividing by x2y1and inserting the power series for gwe obtain u00‡2y0 1 y1‡/C1030 x‡ u0ˆ0: Here and in the following the dots designate terms which are constants or involvepositive powers of x. Now from Eq. (2.30) it follows that y0 1 y1ˆxrÿ1‰ra0‡…r‡1†a1x‡ Š xr‰a0‡a1x‡ Šˆ1 xra0‡…r‡1†a1x‡ a0‡a1x‡ˆr x‡ : 90ORDINARY DIFFERENTIAL EQUATIONS Hence the last equation can be written u00‡2r‡/C1030 x‡ u0ˆ0: …2:32† Since rˆ…1ÿ/C1030†=2 the term …2r‡/C1030†=xequals 1 =x, and by dividing by u0we thus have u00 u0ˆÿ1 x‡ : By integration we obtain lnu0ˆÿlnx‡ oru0ˆ1 xe…...†: Expanding the exponential function in powers of xand integrating once more, we see that the expression for uwill be of the form uˆlnx‡k1x‡k2x2‡ : By inserting this into Eq. (2.31) we find that the second solution is of the form y2…x†ˆy1…x†lnx‡xrX1 mˆ1Amxm: …2:33† /C67ase 3 /C82oots di/C128ering by an integer If the roots r1andr2of the indicial equation (2.29) di/C128er by an integer, say, r1ˆr andr2ˆrÿ/C112, where pis a positive integer, then we may always determine one solution as before, namely, the solution corresponding to r1: y1…x†ˆxr1…a0‡a1x‡a2x2‡ † : To determine a second solution y2, we may proceed as in Case 2. The first steps are literally the same and yield Eq. (2.32). We determine 2 r‡/C1030in Eq. (2.32). Then from the indicial equation (2.29), we find ÿ…r1‡r2†ˆ/C1030ÿ1. In our case, r1ˆrand r2ˆrÿ/C112, therefore, /C1030ÿ1ˆ/C112ÿ2r. Hence in Eq. (2.32) we have 2r‡/C1030ˆ/C112‡1, and we thus obtain u00 u0ˆÿ/C112‡1 x‡ : Integrating, we find lnu0ˆÿ … /C112‡1†lnx‡ oru0ˆxÿ…/C112‡1†e…...†; where the dots stand for some series of positive powers of x. By expanding the exponential function as before we obtain a series of the form u0ˆ1 x/C112‡1‡k1 x/C112‡‡k/C112 x‡k/C112‡1‡k/C112‡2x‡ : 91SOLUTIONS IN POWER SERIES Integrating, we have uˆÿ1 /C112x/C112ÿ‡ k/C112lnx‡k/C112‡1x‡ : …2:34† Multiplying this expression by the series y1…x†ˆxr1…a0‡a1x‡a2x2‡ † and remembering that r1ÿ/C112ˆr2we see that y2ˆuy1is of the form y2…x†ˆk/C112y1…x†lnx‡xr2X1 mˆ0amxm: …2:35† While for a double root of Eq. (2.29) the second solution always contains a logarithmic term, the coecient k/C112may be zero and so the logarithmic term may be missing, as shown by the following example. Example 2.15 Solve the di/C128erential equation x2y00‡xy0‡…x2ÿ1 4†yˆ0: Solution: Substituting Eq. (2.28) and its derivatives into this equation, we obtain X1 mˆ0‰…m‡r†…m‡rÿ1†‡…m‡r†ÿ14Šamxm‡r‡P1 mˆ0amxm‡r‡2ˆ0: By equating the coecient of xrto zero we get the indicial equation r…rÿ1†‡rÿ14ˆ0o r r2ˆ14: The roots r1ˆ12andr2ˆÿ12di/C128er by an integer. By equating the sum of the coecients of xs‡rto zero we find ‰…r‡1†r‡…rÿ1†ÿ14Ša1ˆ0…sˆ1†: …2:36a† ‰…s‡r†…s‡rÿ1†‡s‡rÿ1 4Šas‡asÿ2ˆ0…sˆ2;3;...†: …2:36b† For rˆr1ˆ12, Eq. (2.36a) yields a1ˆ0, and the indicial equation (2.36b) becomes …s‡1†sas‡asÿ2ˆ0: From this and a1ˆ0 we obtain a3ˆ0,a5ˆ0, etc. Solving the indicial equation forasand setting sˆ2/C112, we get a2/C112ˆÿa2/C112ÿ2 2/C112…2/C112‡1†…/C112ˆ1;2;...†: 92ORDINARY DIFFERENTIAL EQUATIONS Hence the non-zero coecients are a2ˆÿa0 3/C33;a4ˆÿa2 45ˆa0 5/C33;a6ˆÿa0 7/C33;etc:; and the solution y1is y1…x†ˆa0xpX1 mˆ0…ÿ1†mx2m …2m‡1†/C33ˆa0xÿ1=2X1 mˆ0…ÿ1†mx2m‡1 …2m‡1†/C33ˆa0sinxxp: …2:37† From Eq. (2.35) we see that a second independent solution is of the form y2…x†ˆky1…x†lnx‡xÿ1=2X1 mˆ0amxm: Substituting this and the derivatives into the di/C128erential equation, we see that the three expressions involving ln xand the expressions ky1andÿky1drop out. Simplifying the remaining equation, we thus obtain 2kxy0 1‡X1 mˆ0m…mÿ1†amxmÿ1=2‡X1 mˆ0amxm‡3=2ˆ0: From Eq. (2.37) we find 2 kxy0ˆÿka0x1=2‡ . Since there is no further term involving x1=2anda06ˆ0, we must have kˆ0. The sum of the coecients of the power xsÿ1=2is s…sÿ1†as‡asÿ2…sˆ2;3;...†: Equating this to zero and solving for as, we have asˆÿasÿ2=‰s…sÿ1†Š … sˆ2;3;...†; from which we obtain a2ˆÿa0 2/C33;a4ˆÿa2 43ˆa0 4/C33;a6ˆÿa0 6/C33;etc:; a3ˆÿa1 3/C33;a5ˆÿa3 54ˆa1 5/C33;a7ˆÿa1 7/C33;etc: We may take a1ˆ0, because the odd powers would yield a1y1=a0. Then y2…x†ˆa0xÿ1=2X1 mˆ0…ÿ1†mx2m …2m†/C33ˆa0cosxxp: /C83imultaneous equations In some physics and engineering problems we may face simultaneous dif- ferential equations in two or more dependent variables. The general solution 93SIMULTANEOUS EQUATIONS of simultaneous equations may be found by solving for each dependent variable separately, as shown by the following example Dx‡2y‡3xˆ0 3x‡Dyÿ2yˆ0) …Dˆd=dt† which can be rewritten as …D‡3†x‡2yˆ0; 3x‡…Dÿ2†yˆ0:) We then operate on the first equation with ( Dÿ2) and multiply the second by a factor 2: …Dÿ2†…D‡3†x‡2…Dÿ2†yˆ0; 6x‡2…Dÿ2†yˆ0:) Subtracting the first from the second leads to …D2‡Dÿ6†xÿ6xˆ…D2‡Dÿ12†xˆ0; which can easily be solved and its solution is of the form x…t†ˆAe3t‡Beÿ4t: Now inserting x…t†back into the original equation to find ygives: y…t†ˆÿ 3Ae3t‡1 2Beÿ4t: /C84he gamma and beta functions The factorial notation n/C33ˆn…nÿ1†…nÿ2† 321 has proved useful in writing down the coecients in some of the series solutions of the di/C128erential equations. However, this notation is meaningless when nis not a positive integer. A useful extension is provided by the gamma (or Euler) function, which is definedby the integral ÿ… †ˆZ 1 0eÿxx ÿ1dx… /C620†… 2:38† and it follows immediately that ÿ…1†ˆZ1 0eÿxdxˆ‰ ÿ eÿxŠ1 0ˆ1: …2:39† Integration by parts gives ÿ… ‡1†ˆZ1 0eÿxx dxˆ‰ ÿ eÿxx Š10‡ Z1 0eÿxx ÿ1dxˆ ÿ… †:…2:40† 94ORDINARY DIFFERENTIAL EQUATIONS When ˆn, a positive integer, repeated application of Eq. (2.40) and use of Eq. (2.39) gives ÿ…n‡1†ˆnÿ…n†ˆn…nÿ1†ÿ…nÿ1†ˆ ...ˆn…nÿ1† 32ÿ…1† ˆn…nÿ1† 321ˆn/C33: Thus the gamma function is a generalization of the factorial function. Eq. (2.40) enables the values of the gamma function for any positive value of to be calculated: thus ÿ…7 2†ˆ…52†ÿ…52†ˆ…52†…32†ÿ…32†ˆ…52†…32†…12†ÿ…12†: Write uˆ‡xpin Eq. (2.38) and we then obtain ÿ… †ˆ2Z1 0u2 ÿ1eÿu2du; so that ÿ…1 2†ˆ2Z1 0eÿu2duˆp: The function ÿ… †has been tabulated for values of between 0 and 1. When <0 we can define ÿ… †with the help of Eq. (2.40) and write ÿ… †ˆ ÿ… ‡1†= : Thus ÿ…ÿ3 2†ˆÿ23ÿ…ÿ12†ˆÿ23…ÿ21†ÿ…12†ˆ43p: When !0;R1 0eÿxx ÿ1dxdiverges so that ÿ…0†is not defined. Another function which will be useful later is the beta function which is defined by B…/C112;/C113†ˆZ1 0t/C112ÿ1…1ÿt†/C113ÿ1dt…/C112;/C113/C620†: …2:41† Substituting tˆ/C118=…1‡/C118†, this can be written in the alternative form B…/C112;/C113†ˆZ1 0/C118/C112ÿ1…1‡/C118†ÿ/C112ÿ/C113d/C118: …2:42† By writing t0ˆ1ÿtwe deduce that B…/C112;/C113†ˆB…/C113;/C112†. The beta function can be expressed in terms of gamma functions as follows: B…/C112;/C113†ˆÿ…/C112†ÿ…/C113† ÿ…/C112‡/C113†: …2:43† To prove this, write xˆat(a/C620) in the integral (2.38) defining ÿ… †, and it is straightforward to show that ÿ… † a ˆZ1 0eÿatt ÿ1dt …2:44† 95THE GAMMA AND BETA FUNCTIONS and, with ˆ/C112‡/C113,aˆ1‡/C118, this can be written ÿ…/C112‡/C113†…1‡/C118†ÿ/C112ÿ/C113ˆZ1 0eÿ…1‡/C118†tt/C112‡/C113ÿ1dt: Multiplying by /C118/C112ÿ1and integrating with respect to /C118between 0 and 1, ÿ…/C112‡/C113†Z1 0/C118/C112ÿ1…1‡/C118†ÿ/C112ÿ/C113d/C118ˆZ1 0/C118/C112ÿ1d/C118Z1 0eÿ…1‡/C118†tt/C112‡/C113‡1dt: Then interchanging the order of integration in the double integral on the right and using Eq. (2.42), ÿ…/C112‡/C113†B…/C112;/C113†ˆZ1 0eÿtt/C112‡/C113ÿ1dtZ1 0eÿ/C118t/C118/C112ÿ1d/C118 ˆZ1 0eÿtt/C112‡/C113ÿ1ÿ…/C112† t/C112dt;using Eq :…2:44† ˆÿ…/C112†Z1 0eÿtt/C113ÿ1dtˆÿ…/C112†ÿ…/C113†: Example 2.15 Evaluate the integralZ1 03ÿ4x2dx: Solution: We first notice that 3 ˆeln 3, so we can rewrite the integral as Z1 03ÿ4x2dxˆZ1 0…eln 3†…ÿ4x2†dxˆZ1 0eÿ…4l n3†x2dx: Now let (4 ln 3) x2ˆz, then the integral becomes Z1 0eÿzdz1=2  4l n3p/C32! ˆ1 24l n3pZ 1 0zÿ1=2eÿzdzˆÿ…1 2† 2 4l n3p ˆp 2 4l n3p : Problems 2.1 Solve the following equations: (a)xdy=dx‡y2ˆ1; (b)dy=dxˆ…x‡y†2. 2.2 Melting of a sphere of ice: Assume that a sphere of ice melts at a rate proportional to its surface area. Find an expression for the volume at any time t. 2.3 Show that …3x2‡ycosx†dx‡…sinxÿ4y3†dyˆ0 is an exact di/C128erential equation and find its general solution. 96ORDINARY DIFFERENTIAL EQUATIONS 2.4 /C82/C67circuits: A typical /C82/C67circuit is shown in Fig. 2.3. Find current flow I…t† in the circuit, assuming /C69…t†ˆ/C690. Hint: the voltage drop across the capacitor is given /C81/C47/C67, with /C81(t) the charge on the capacitor at time t. 2.5 Find a constant such that …x‡y† is an integrating factor of the equation …4x2‡2xy‡6y†dx‡…2x2‡9y‡3x†dyˆ0: What is the solution of this equation/C63 2.6 Solve dy=dx‡yˆy3x: 2.7 Solve: (a) the equation …D2ÿDÿ12†yˆ0 with the boundary conditions yˆ0, Dyˆ3 when tˆ0; (b) the equation …D2‡2D‡3†yˆ0 with the boundary conditions yˆ2, Dyˆ0 when tˆ0; (c) the equation …D2ÿ2D‡1†yˆ0 with the boundary conditions yˆ5, Dyˆ3 when tˆ0. 2.8 Find the particular integral of …D2‡2Dÿ1†yˆ3‡t3. 2.9 Find the particular integral of …2D2‡5D‡7†ˆ3e2t. 2.10 Find the particular integral of …3D2‡Dÿ5†yˆcos 3 t: 2.11 Simple harmonic motion of a pendulum (Fig. 2.4): Suspend a ball of mass m at the end of a massless rod of length Land set it in motion swinging back and forth in a vertical plane. Show that the equation of motion of the ball is d2 dt2‡/C103 Lsinˆ0; where gis the local gravitational acceleration. Solve this pendulum equation for small displacements by replacing sin by. 2.12 Forced oscillations with damping: If we allow an external driving force /C70…t† in addition to damping (Example 2.12), the motion of the oscillator is governed by y00‡b my0‡k myˆ/C70…t†; 97PROBLEMS Figure 2.3. /C82/C67circuit. a constant coecient non-homogeneous equation. Solve this equation for /C70…t†ˆAcos…/C33t†: 2.13 Solve the equation r2d2R dr2‡2rdR drÿn…n‡1†Rˆ0…nconstant †: 2.14 The first-order non-linear equation dy dx‡y2‡/C81…x†y‡R…x†ˆ0 is known as Riccati’s equation. Show that, by use of a change of dependentvariable yˆ1 zdz dx; Riccati’s equation transforms into a second-order linear di/C128erential equa-tion d2z dx2‡/C81…x†dz dx‡R…x†zˆ0: Sometimes Riccati’s equation is written as dy dx‡P…x†y2‡/C81…x†y‡R…x†ˆ0: Then the transformation becomes yˆÿ1 P…x†dz dx and the second-order equation takes the form d2z dx2‡/C81‡1 PdP dxdz dx‡PRzˆ0: 98ORDINARY DIFFERENTIAL EQUATIONS Figure 2.4. Simple pendulum. 2.15 Solve the equation 4 x2y00‡4xy0‡…x2ÿ1†yˆ0 by using Frobenius’ method, where y0ˆdy=dx,a n d y00ˆd2y=dx2. 2.16 Find a series solution, valid for large values of x, of the equation …1ÿx2†y00ÿ2xy0‡2yˆ0: 2.17 Show that a series solution of Airy’s equation y00ÿxyˆ0i s yˆa01‡x3 23‡x6 2356‡/C32! ‡b0x‡x4 34‡x7 3467‡/C32! : 2.18 Show that Weber’s equation y00‡…n‡1 2ÿ14x2†yˆ0 is reduced by the sub- stitution yˆeÿx2=4/C118to the equation d2/C118=dx2ÿx…d/C118=dx†‡n/C118ˆ0. Show that two solutions of this latter equation are /C1181ˆ1ÿn 2/C33x2‡n…nÿ2† 4/C33x4ÿn…nÿ2†…nÿ4† 6/C33x6‡ÿ ; /C1182ˆxÿ…nÿ1† 3/C33x3‡…nÿ1†…nÿ3† 5/C33x5ÿ…nÿ1†…nÿ3†…nÿ5† 7/C33x7‡ÿ : 2.19 Solve the following simultaneous equations Dx‡yˆt3 Dyÿxˆt) …Dˆd=dt†: 2.20 Evaluate the integrals: (a)Z1 0x3eÿxdx: (b)Z1 0x6eÿ2xdx…hint: let yˆ2x†: (c)Z1 0ypeÿy2dy…hint: let y2ˆx†: (d)Z1 0dx ÿlnxp …hint : let ÿlnxˆu†: 2.21 ( a) Prove that B…/C112;/C113†ˆ2Z=2 0sin2/C112ÿ1cos2/C113ÿ1d. (b) Evaluate the integralZ1 0x4…1ÿx†3dx: 2.22 Show that n/C33/C25 2np nneÿn. This is known as Stirling’s factorial approxima- tion or asymptotic formula for n/C33. 99PROBLEMS 3 Matrix algebra As vector methods have become standard tools for physicists, so too matrix methods are becoming very useful tools in sciences and engineering. Matrices occur in physics in at least two ways: in handling the eigenvalue problems in classical and quantum mechanics, and in the solutions of systems of linear equa- tions. In this chapter, we introduce matrices and related concepts, and define some basic matrix algebra. In Chapter 5 we will discuss various operations with matrices in dealing with transformations of vectors in vector spaces and the operation of linear operators on vector spaces. /C68efinition of a matri/C120 A matrix consists of a rectangular block or ordered array of numbers that obeys prescribed rules of addition and multiplication. The numbers may be real or complex. The array is usually enclosed within curved brackets. Thus 12 4 2ÿ17 is a matrix consisting of 2 rows and 3 columns, and it is called a 2 3( 2b y3 ) matrix. An mnmatrix consists of mrows and ncolumns, which is usually expressed in a double sux notation: ~Aˆa11a12a13 a1n a21a22a23 ... a2n ............ am1am2am3 ...amn0 BBBB@1 CCCCA: …3:1† Each number a ijis called an element of the matrix, where the first subscript i denotes the row, while the second subscript jindicates the column. Thus, a23 100 refers to the element in the second row and third column. The element aijshould be distinguished from the element aji. It should be pointed out that a matrix has no single numerical value; therefore it must be carefully distinguished from a determinant. We will denote a matrix by a letter with a tilde over it, such as ~Ain (3.1). Sometimes we write ( aij)o r( aij†mn, if we wish to express explicitly the particular form of element contained in ~A. Although we have defined a matrix here with reference to numbers, it is easy to extend the definition to a matrix whose elements are functions fi…x†; for a 2 3 matrix, for example, we have f1…x†f2…x†f3…x† f4…x†f5…x†f6…x† : A matrix having only one row is called a row matrix or a row vector, while a matrix having only one column is called a column matrix or a column vector. An ordinary vector /C65ˆA1^e1‡A2^e2‡A3^e3can be represented either by a row matrix or by a column matrix. If the numbers of rows mand columns nare equal, the matrix is called a square matrix of order n. In a square matrix of order n, the elements a11;a22;...;annform what is called the principal (or leading) diagonal, that is, the diagonal from the top left hand corner to the bottom right hand corner. The diagonal from the top right hand corner to the bottom left hand corner is sometimes termed the trailing diagonal. Only a square matrix possesses a principal diagonal and a trailing diagonal. The sum of all elements down the principal diagonal is called the trace, or spur, of the matrix. We write Tr ~AˆXn iˆ1aii: If all elements of the principal diagonal of a square matrix are unity while all other elements are zero, then it is called a unit matrix (for a reason to be explained later) and is denoted by ~I. Thus the unit matrix of order 3 is ~Iˆ100 010 0010 B@1 CA: A square matrix in which all elements other than those along the principal diagonal are zero is called a diagonal matrix. A matrix with all elements zero is known as the null (or zero) matrix and is denoted by the symbol ~0, since it is not an ordinary number, but an array of zeros. 101DEFINITION OF A MATRI/C88 Four basic algebra operations for matrices Equality of matrices Two matrices ~Aˆ…ajk†and ~Bˆ…bjk†are equal if and only if ~Aand ~Bhave the same order (equal numbers of rows and columns) and corresponding elements are equal, that is ajkˆbjkfor all jandk: Then we write ~Aˆ~B: /C65ddition of matrices Addition of matrices is defined only for matrices of the same order. If ~Aˆ…ajk† and ~Bˆ…bjk†have the same order, the sum of ~Aand ~Bis a matrix of the same order ~Cˆ~A‡~B with elements cjkˆajk‡bjk: …3:2† We see that ~Cis obtained by adding corresponding elements of ~Aand ~B. Example 3.1 If ~Aˆ214 302 ; ~Bˆ35 121 ÿ3 hen ~Cˆ~A‡~Bˆ214 302/C32! ‡35 121 ÿ3/C32! ˆ2‡31‡54‡1 3‡20‡12ÿ3/C32! ˆ56 551 ÿ1/C32! : From the definitions we see that matrix addition obeys the commutative and associative laws, that is, for any matrices ~A,~B,~Cof the same order ~A‡~Bˆ~B‡~A; ~A‡… ~B‡~C†ˆ… ~A‡~B†‡ ~C: …3:3† Similarly, if ~Aˆ…a jk†and ~Bˆ…bjk) have the same order, we define the di/C128er- ence of ~Aand ~Bas ~Dˆ~Aÿ~B 102MATRI/C88 ALGEBRA with elements djkˆajkÿbjk: …3:4† /C77ultiplication of a matrix by a number If~Aˆ…ajk†andcis a number (or scalar), then we define the product of ~Aandcas c~Aˆ~Acˆ…cajk†; …3:5† we see that c~Ais the matrix obtained by multiplying each element of ~Abyc. We see from the definition that for any matrices and any numbers, c…~A‡~B†ˆc~A‡c~B;…c‡k†~Aˆc~A‡k~A;c…k~A†ˆck~A: …3:6† Example 3.2 7abc def ˆ7a7b7c 7d7e7f : Formulas (3.3) and (3.6) express the properties which are characteristic for a vector space. This gives vector spaces of matrices. We will discuss this further in Chapter 5. /C77atrix multiplication The matrix product ~A~Bof the matrices ~Aand ~Bis defined if and only if the number of columns in ~Ais equal to the number of rows in ~B. Such matrices are sometimes called ‘conformable’. If ~Aˆ…ajk†is an nsmatrix and ~Bˆ…bjk†is an smmatrix, then ~Aand ~Bare conformable and their matrix product, written ~Cˆ~A~B,i sa n nmmatrix formed according to the rule cikˆXs jˆ1aijbjk;iˆ1;2;...;nk ˆ1;2;...;m: …3:7† Consequently, to determine the ijth element of matrix ~C, the corresponding terms of the ith row of ~Aandjth column of ~Bare multiplied and the resulting products added to form cij. Example 3.3Let ~Aˆ21 4 ÿ302 ; ~Bˆ35 2ÿ1 420 B@1 CA 103FOUR BASIC ALGEBRA OPERATIONS FOR MATRICES then ~A~Bˆ23‡12‡442 5‡1… ÿ 1†‡42 …ÿ3†3‡02‡24…ÿ3†5‡0… ÿ 1†‡22/C32! ˆ24 17 ÿ1ÿ11/C32! : The reader should master matrix multiplication, since it is used throughout the rest of the book. In general, matrix multiplication is not commutative: ~A~B6ˆ~B~A. In fact, ~B~Ais often not defined for non-square matrices, as shown in the following example. Example 3.4 If ~Aˆ12 34/C32! ; ~Bˆ37/C32! then ~A~Bˆ12 3437 ˆ13‡27 33‡47 ˆ1737 : But ~B~Aˆ371234 is not defined. Matrix multiplication is associative and distributive: …~A~B†~Cˆ~A…~B~C†;…~A‡~B†~Cˆ~A~C‡~B~C: To prove the associative law, we start with the matrix product ~A~B, then multi- ply this product from the right by ~C: ~A~BˆX kaikbkj; …~A~B†~CˆX jX kaikbkj/C32! cjs"# ˆX kaikX jbkjcjs/C32! ˆ~A…~B~C†: Products of matrices di/C128er from products of ordinary numbers in many remarkable ways. For example, ~A~Bˆ0 does not imply ~Aˆ0o r ~Bˆ0. Even more bizarre is the case where ~A2ˆ0, ~A6ˆ0; an example of which is ~Aˆ0100 : 104MATRI/C88 ALGEBRA When you first run into Eq. (3.7), the rule for matrix multiplication, you might ask how anyone would arrive at it. It is suggested by the use of matrices in connection with linear transformations. For simplicity, we consider a very simple case: three coordinates systems in the plane denoted by the x1x2-system, the y1y2- system, and the z1z2-system. We assume that these systems are related by the following linear transformations x1ˆa11y1‡a12y2;x2ˆa21y1‡a22y2; …3:8† y1ˆb11z1‡b12z2;y2ˆb21z1‡b22z2: …3:9† Clearly, the x1x2-coordinates can be obtained directly from the z1z2-coordinates by a single linear transformation x1ˆc11z1‡c12z2;x2ˆc21z1‡c22z2; …3:10† whose coecients can be found by inserting (3.9) into (3.8), x1ˆa11…b11z1‡b12z2†‡a12…b21z1‡b22z2†; x2ˆa21…b11z1‡b12z2†‡a22…b21z1‡b22z2†: Comparing this with (3.10), we find c11ˆa11b11‡a12b21;c12ˆa11b12‡a12b22; c21ˆa21b11‡a22b21;c22ˆa21b12‡a22b22; or briefly cjkˆX2 iˆ1ajibik;j;kˆ1;2; …3:11† which is in the form of (3.7). Now we rewrite the transformations (3.8), (3.9) and (3.10) in matrix form: ~Xˆ~A~Y; ~Yˆ~B~Z;and ~Xˆ~C~Z; where ~Xˆx1 x2/C32! ; ~Yˆy1 y2/C32! ; ~Zˆz1 z2/C32! ; ~Aˆa11a12 a21a22/C32! ; ~Bˆb11b12 b21b22/C32! ;C::: ˆc11c12 c21c22/C32! : We then see that ~Cˆ~A~B, and the elements of ~Care given by (3.11). Example 3.5 Rotations in three-dimensional space: An example of the use of matrix multi- plication is provided by the representation of rotations in three-dimensional 105FOUR BASIC ALGEBRA OPERATIONS FOR MATRICES space. In Fig. 3.1, the primed coordinates are obtained from the unprimed coor- dinates by a rotation through an angle about the x3-axis. We see that x0 1is the sum of the projection of x1onto the x0 1-axis and the projection of x2onto the x0 1- axis: x0 1ˆx1cos‡x2cos…=2ÿ†ˆx1cos‡x2sin; similarly x0 2ˆx1cos…=2‡†‡x2cosˆÿx1sin‡x2cos and x0 3ˆx3: We can put these in matrix form X0ˆRX; where X0ˆx0 1 x0 2 x0 30 B@1 CA;Xˆx1 x2 x30 B@1 CA;Rˆcossin0 ÿsincos0 00 10 B@1 CA: 106MATRI/C88 ALGEBRA Figure 3.1. Coordinate changes by rotation. /C84he commutator Even if matrices ~Aand ~Bare both square matrices of order n, the products ~A~B and ~B~A, although both square matrices of order n, are in general quite di/C128erent, since their individual elements are formed di/C128erently. For example, 12 131012 ˆ3446 but10121213 ˆ1238 : The di/C128erence between the two products ~A~Band ~B~Ais known as the commu- tator of ~Aand ~Band is denoted by ‰~A;~BŠˆ ~A~Bÿ~B~A: …3:12† It is obvious that ‰~B;~AŠˆÿ ‰ ~A;~BŠ: …3:13† If two square matrices ~Aand ~Bare very carefully chosen, it is possible to make the product identical. That is ~A~Bˆ~B~A. Two such matrices are said to commute with each other. Commuting matrices play an important role in quantum mechanics. If~Acommutes with ~Band ~Bcommutes with ~C, it does not necessarily follow that ~Acommutes with ~C. Po/C119ers of a matri/C120 Ifnis a positive integer and ~Ais a square matrix, then ~A 2ˆ~A~A,~A3ˆ~A~A~A, and in general, ~Anˆ~A~A ~A(ntimes). In particular, ~A0ˆ~I. Functions of matrices As we define and study various functions of a variable in algebra, it is possible to define and evaluate functions of matrices. We shall briefly discuss the following functions of matrices in this section: integral powers and exponential. A simple example of integral powers of a matrix is polynomials such as f…~A†ˆ ~A2‡3~A5: Note that a matrix can be multiplied by itself if and only if it is a square matrix.Thus ~Ahere is a square matrix and we denote the product ~A~Aas ~A 2. More fancy examples can be obtained by taking series, such as ~SˆX1 kˆ0ak~Ak; where akare scalar coecients. Of course, the sum has no meaning if it does not converge. The convergence of the matrix series means every matrix element of the 107THE COMMUTATOR infinite sum of matrices converges to a limit. We will not discuss the general theory of convergence of matrix functions. Another very common series is definedby e~AˆX1 nˆ0~An n/C33: /C84ranspose of a matri/C120 Consider an mnmatrix ~A, if the rows and columns are systematically changed to columns to rows, without changing the order in which they occur, the newmatrix is called the transpose of matrix ~A. It is denoted by ~A T: ~Aˆa11a12a13 ... a1n a21a22a23 ... a2n ............ am1am2am3 ...amn0 BBBB@1 CCCCA; ~A Tˆa11a21a31 ...am1 a12a22a32 ...am2 ............ an1a2na3n ...amn0 BBBB@1 CCCCA: Thus the transpose matrix has nrows and mcolumns. If ~Ais written as ( a jk), then ~ATmay be written as …akj). ~Aˆ…ajk†; ~ATˆ…akj†: …3:14† The transpose of a row matrix is a column matrix, and vice versa. Example 3.6 ~Aˆ123 456 ; ~ATˆ14 25 360 B@1 CA; ~Bˆ…123 †; ~BTˆ1 2 30 B@1 CA: It is obvious that …~AT†Tˆ~A,a n d …~A‡~B†Tˆ~AT‡~BT. It is also easy to prove that the transpose of the product is the product of the transposes in reverse: …~A~B†Tˆ~BT~AT: …3:15† Proof: …~A~B†T ijˆ… ~A~B†jiby definition ˆX kAjkBki ˆX kBT ikATkj ˆ… ~BT~AT†ij 108MATRI/C88 ALGEBRA so that …AB†TˆBTATq:e:d: Because of (3.15), even if ~Aˆ~ATand ~Bˆ~BT,…~A~B†T6ˆ~A~Bunless the matrices commute. /C83/C121mmetric and s/C107e/C119-s/C121mmetric matrices A square matrix ~Aˆ…ajk†is said to be symmetric if all its elements satisfy the equations akjˆajk; …3:16† that is, ~Aand its transpose are equal ~Aˆ~AT. For example, ~Aˆ157 53 ÿ4 7ÿ400 B@1 CA is a third-order symmetric matrix: the elements of the ith row equal the elements ofith column, for all i. On the other hand, if the elements of ~Asatisfy the equations akjˆÿajk; …3:17† then ~Ais said to be skew-symmetric, or antisymmetric. Thus, for a skew-sym- metric ~A, its transpose equals minus ÿ~A:~ATˆÿ ~A. Since the elements ajjalong the principal diagonal satisfy the equations ajjˆÿajj, it is evident that they must all vanish. For example, ~Aˆ0ÿ25 20 1 ÿ5ÿ100 B@1 CA is a skew-symmetric matrix. Any real square matrix ~Amay be expressed as the sum of a symmetric matrix ~R and a skew-symmetric matrix ~S, where ~Rˆ1 2…~A‡~AT†and ~Sˆ12…~Aÿ~AT†: …3:18† Example 3.7 The matrix ~Aˆ23 5ÿ1 may be written in the form ~Aˆ~R‡~S, where ~Rˆ1 2…~A‡~AT†ˆ24 4ÿ1 ~Sˆ1 2…~Aÿ~AT†ˆ0ÿ1 10 : 109SYMMETRIC AND SKEW-SYMMETRIC MATRICES The product of two symmetric matrices need not be symmetric. This is because of (3.15): even if ~Aˆ~ATand ~Bˆ~BT,…~A~B†T6ˆ~A~Bunless the matrices commute. A square matrix whose elements above or below the principal diagonal are all zero is called a triangular matrix. The following two matrices are triangular matrices: 100 230 5020 B@1 CA;16 ÿ1 02 3 00 40 B@1 CA: A square matrix ~Ais said to be singular if det ~Aˆ0, and non-singular if det ~A6ˆ0, where det ~Ais the determinant of the matrix ~A. /C84he matri/C120 representation of a /C118ector product The scalar product defined in ordinary vector theory has its counterpart in matrix theory. Consider two vectors /C65ˆ…A1;A2;A3†and/C66ˆ…B1;B2;B3†the counter- part of the scalar product is given by ~A~BTˆ…A1A2A3†B1 B2 B30 B@1 CAˆA1B1‡A2B2‡A3B3: Note that ~B~ATis the transpose of ~A~BT,a n d ,b e i n ga1 1 matrix, the transpose equals itself. Thus a scalar product may be written in these two equivalent forms. Similarly, the vector product used in ordinary vector theory must be replaced by something more in keeping with the definition of matrix multiplication. Note that the vector product /C65/C66ˆ…A2B3ÿA3B2†^e1‡…A3B1ÿA1B3†^e2‡…A1B2ÿA2B1†^e3 can be represented by the column matrix A2B3ÿA3B2 A3B1ÿA1B3 A1B2ÿA2B10 B@1 CA: This can be split into the product of two matrices A2B3ÿA3B2 A3B1ÿA1B3 A1B2ÿA2B10 B@1 CAˆ0ÿA2A2 A3 0ÿA1 ÿA2A1 00 B@1 CAB1 B2 B30 B@1 CA 110MATRI/C88 ALGEBRA or A2B3ÿA3B2 A3B1ÿA1B3 A1B2ÿA2B10 B@1 CAˆ0ÿB2B2 B3 0ÿB1 ÿB2B1 00 B@1 CAA1 A2 A30 B@1 CA: Thus the vector product may be represented as the product of a skew-symmetric matrix and a column matrix. However, this definition only holds for 3 3 matrices. Similarly, curl Amay be represented in terms of a skew-symmetric matrix operator, given in Cartesian coordinates by /C114Aˆ0 ÿ/C64=/C64x3/C64=/C64x2 /C64=/C64x3 0 ÿ/C64=/C64x1 ÿ/C64=/C64x2/C64=/C64x1 00 B@1 CAA1 A2 A30 B@1 CA: In a similar way, we can investigate the triple scalar product and the triple vector product. /C84he in/C118erse of a matri/C120 If for a given square matrix ~Athere exists a matrix ~Bsuch that ~A~Bˆ~B~Aˆ~I, where ~Iis a unit matrix, then ~Bis called an inverse of matrix ~A. Example 3.8The matrix ~Bˆ35 12 is an inverse of ~Aˆ2ÿ5 ÿ13 ; since ~A~Bˆ2ÿ5 ÿ1335 12 ˆ1001 ˆ~I and ~B~Aˆ35122ÿ5 ÿ13 ˆ1001 ˆ~I: An invertible matrix has a unique inverse. That is, if ~Band ~Care both inverses of the matrix ~A, then ~Bˆ~C. The proof is simple. Since ~Bis an inverse of ~A, 111THE INVERSE OF A MATRI/C88 ~B~Aˆ~I. Multiplying both sides on the right by ~Cgives …~/C66~A†~Cˆ~I~Cˆ~C. On the other hand, ( ~B~A†~Cˆ~B…~A~C†ˆ ~B~Iˆ~B, so that ~Bˆ~C. As a consequence of this result, we can now speak of theinverse of an invertible matrix. If ~Ais invertible, then its inverse will be denoted by ~Aÿ1. Thus ~A~Aÿ1ˆ~Aÿ1~Aˆ~I: …3:19† It is obvious that the inverse of the inverse is the given matrix, that is, …~Aÿ1†ÿ1ˆ~A: …3:20† It is easy to prove that the inverse of the product is the product of the inverse in reverse order, that is, …~A~B†ÿ1ˆ~Bÿ1~Aÿ1: …3:21† To prove (3.21), we start with ~A~Aÿ1ˆ~I, with ~Areplaced by ~A~B, that is, ~A~B…~A~B†ÿ1ˆ~I: By premultiplying this by ~Aÿ1we get ~B…~A~B†ÿ1ˆ~Aÿ1: If we premultiply this by ~Bÿ1, the result follows. /C65 method for finding ~Aÿ1 The positive power for a square matrix ~Ais defined as ~Anˆ~A~A ~A(nfactors) and ~A0ˆ~I, where nis a positive integer. If, in addition, ~Ais invertible, we define ~Aÿnˆ… ~Aÿ1†nˆ~Aÿ1~Aÿ1 ~Aÿ1…nfactors †: We are now in position to construct the inverse of an invertible matrix ~A: ~Aˆa11a12 ...a1n a21a22 ...a2n ......... an1an2 ...ann0 BBBB@1 CCCCA: Thea jkare known. Now let ~Aÿ1ˆa0 11a0 12 ...a0 1n a0 21a0 22 ...a0 2n ......... a0 n1a0 n2 ...a0 nn0 BBBB@1 CCCCA: 112MATRI/C88 ALGEBRA Thea0 jkare required to construct ~Aÿ1. Since ~A~Aÿ1ˆ~I, we have a11a0 11‡a12a0 12‡‡ a1na0 1nˆ1; a21a0 21‡a22a0 22‡‡ a2na0 2nˆ0; ... an1a0 n1‡an2a0 n2‡‡ anna0 nnˆ0:…3:22† The solution to the above set of linear algebraic equations (3.22) may be facili- tated by applying Cramer’s rule. Thus a0 jkˆcofactor akj det ~A: …3:23† From (3.23) it is clear that ~Aÿ1exists if and only if matrix ~Ais non-singular (that is, det ~A6ˆ0). /C83/C121stems of linear equations and the in/C118erse of a matri/C120 As an immediate application, let us apply the concept of an inverse matrix to a system of nlinear equations in nunknowns …x1;...;xn†: a11x1‡a12x2‡‡ a1nxnˆb1; a21x2‡a22x2‡‡ a2nxnˆb2; ... an1xn‡an2xn‡‡ annxnˆbn; in matrix form we have ~A~Xˆ~B; …3:24† where ~Aˆa11a12 ...a1n a21a22 ...a2n ......... an1an2 ...ann0 BBBB@1 CCCCA; ~Xˆx 1 x2 ... xn0 BBBB@1 CCCCA; ~Bˆb 1 b2 ... bn0 BBBB@1 CCCCA: We can prove that the above linear system possesses a unique solution given by ~Xˆ~A ÿ1~B: …3:25† The proof is simple. If ~Ais non-singular it has a unique inverse ~Aÿ1. Now pre- multiplying (3.24) by ~Aÿ1we obtain ~Aÿ1…~A~X†ˆ ~Aÿ1~B; 113SYSTEMS OF LINEAR EQUATIONS but ~Aÿ1…~A~X†ˆ… ~Aÿ1~A†~Xˆ~X so that ~Xˆ~Aÿ1~Bis a solution to …3:24†;~A~Xˆ~B: /C67omple/C120 con/C106ugate of a matri/C120 If~Aˆ…ajk†is an arbitrary matrix whose elements may be complex numbers, the complex conjugate matrix, denoted by ~A/C42, is also a matrix of the same order, every element of which is the complex conjugate of the corresponding element of ~A, that is, …A/C42†jkˆa/C42jk: …3:26† /C72ermitian con/C106ugation If~Aˆ…ajk†is an arbitrary matrix whose elements may be complex numbers, when the two operations of transposition and complex conjugation are carried out on ~A, the resulting matrix is called the hermitian conjugate (or hermitian adjoint) of the original matrix ~Aand will be denoted by ~A/C121. We frequently call ~A/C121A-dagger. The order of the two operations is immaterial: ~A/C121ˆ… ~AT†/C42ˆ… ~A/C42†T: …3:27† In terms of the elements, we have …~A/C121†jkˆa/C42kj: …3:27a† It is clear that if ~Ais a matrix of order mn, then ~A/C121is a matrix of order nm. We can prove that, as in the case of the transpose of a product, the adjoint of the product is the product of the adjoints in reverse: …~A~B†/C121ˆ~B/C121~A/C121: …3:28† /C72ermitian/C47anti-hermitian matri/C120 A matrix ~Athat obeys ~A/C121ˆ~A …3:29† is called a hermitian matrix. It is very clear the following matrices are hermitian: 1ÿi i2 ;45 ‡2i6‡3i 5ÿ2i 5 ÿ1ÿ2i 6ÿ3iÿ1‡2i 60 B@1 CA;where iˆ ÿ1p : 114MATRI/C88 ALGEBRA Evidently all the elements along the principal diagonal of a hermitian matrix must be real. A hermitian matrix is also defined as a matrix whose transpose equals its complex conjugate: ~ATˆ~A/C42…that is ;akjˆa/C42jk†: …3:29a† These two definitions are the same. First note that the elements in the principaldiagonal of a hermitian matrix are always real. Furthermore, any real symmetric matrix is hermitian, so a real hermitian matrix is a symmetric matrix. The product of two hermitian matrices is not generally hermitian unless they commute. This is because of property (3.28): even if ~A /C121ˆ~Aand ~B/C121ˆ~B, …~A~B†/C1216ˆ~A~Bunless the matrices commute. A matrix ~Athat obeys ~A/C121ˆÿ ~A …3:30† is called an anti-hermitian (or skew-hermitian) matrix. All the elements along the principal diagonal must be pure imaginary. An example is 6i 5‡2i6‡3i ÿ5‡2iÿ8iÿ1ÿ2i ÿ6‡3i1ÿ2i 00 B@1 CA: We summarize the three operations on matrices discussed above in Table 3.1. /C79rthogonal matri/C120 (real) A matrix ~Aˆ…ajk†mnsatisfying the relations ~A~ATˆ~In; …3:31a† ~AT~Aˆ~Im …3:31b† is called an orthogonal matrix. It can be shown that if ~Ais a finite matrix satisfy- ing both relations (3.31a) and (3.31b), then ~Amust be square, and we have ~A~ATˆ~AT~Aˆ~I: …3:32† 115ORTHOGONAL MATRI/C88 (REAL) Table 3.1. Operations on matrices Operation Matrix element ~A ~B If~Bˆ~A Transposition ~Bˆ~ATbijˆaji mnn m Symmetrica Complex conjugation ~Bˆ~A/C42bijˆa/C42ij mnm n Real Hermitian conjugation ~Bˆ~AT/C42bijˆa/C42ji mnn m Hermitian aFor square matrices only. But if ~Ais an infinite matrix, then ~Ais orthogonal if and only if both (3.31a) and (3.31b) are simultaneously satisfied. Now taking the determinant of both sides of Eq. (3.32), we have (det ~A†2ˆ1, or det ~Aˆ1. This shows that ~Ais non-singular, and so ~Aÿ1exists. Premultiplying (3.32) by ~Aÿ1we have ~Aÿ1ˆ~AT: …3:33† This is often used as an alternative way of defining an orthogonal matrix. The elements of an orthogonal matrix are not all independent. To find the conditions between them, let us first equate the ijth element of both sides of ~A~ATˆ~I; we find that Xn kˆ1aikajkˆij: …3:34a† Similarly, equating the ijth element of both sides of ~AT~Aˆ~I, we obtain Xn kˆ1akiakjˆij: …3:34b† Note that either (3.34a) and (3.34b) gives 2 n…n‡1†relations. Thus, for a real orthogonal matrix of order n, there are only n2ÿn…n‡1†=2ˆn…nÿ1†=2 di/C128er- ent elements. /C85nitar/C121 matri/C120 A matrix ~Uˆ…ujk†mnsatisfying the relations ~U~U/C121ˆ~In; …3:35a† ~U/C121~Uˆ~Im …3:35b† is called a unitary matrix. If ~Uis a finite matrix satisfying both (3.35a) and (3.35b), then ~Umust be a square matrix, and we have ~U~U/C121ˆ ~U/C121~Uˆ~I: …3:36† This is the complex generalization of the real orthogonal matrix. The elements of a unitary matrix may be complex, for example 1 2p1i i1 is unitary. From the definition (3.35), a real unitary matrix is orthogonal. Taking the determinant of both sides of (3.36) and noting that det ~U/C121ˆ…det ~U)/C42, we have …det ~U†…det ~U†/C42ˆ1o r jdet ~Ujˆ1: …3:37† 116MATRI/C88 ALGEBRA This shows that the determinant of a unitary matrix can be a complex number of unit magnitude, that is, a number of the form ei , where is a real number. It also shows that a unitary matrix is non-singular and possesses an inverse. Premultiplying (3.35a) by ~Uÿ1, we get ~U/C121ˆ~Uÿ1: …3:38† This is often used as an alternative way of defining a unitary matrix. Just as in the case of an orthogonal matrix that is a special (real) case of a unitary matrix, the elements of a unitary matrix satisfy the following conditions: Xn kˆ1uikujk/C42ˆij;Xn kˆ1ukiukj/C42ˆij: …3:39† The product of two unitary matrices is unitary. The reason is as follows. If ~U1 and ~U2are two unitary matrices, then ~U1~U2…~U1~U2†/C121ˆ~U1~U2…~U/C121 2~U/C121 1†ˆ ~U1~U/C121 1ˆ~I; …3:40† which shows that U1U2is unitary. /C82otation matrices Let us revisit Example 3.5. Our discussion will illustrate the power and usefulnessof matrix methods. We will also see that rotation matrices are orthogonal matrices. Consider a point Pwith Cartesian coordinates …x 1;x2;x3†(see Fig. 3.2). We rotate the coordinate axes about the x3-axis through an angle and create a new coordinate system, the primed system. The point Pnow has the coordinates …x0 1;x0 2;x0 3†in the primed system. Thus the position vector rof point Pcan be written as rˆX3 iˆ1xi^eiˆX3 iˆ1x0 i^e0 i: …3:41† 117ROTATION MATRICES Figure 3.2. Coordinate change by rotation. Taking the dot product of Eq. (3.41) with ^e0 1and using the orthonormal relation ^e0 i^e0 jˆij(where ijis the Kronecker delta symbol), we obtain x0 1ˆr^e0 1. Similarly, we have x0 2ˆr^e0 2andx0 3ˆr^e0 3. Combining these results we have x0 iˆX3 jˆ1^e0 i^ejxjˆX3 jˆ1ijxj; iˆ1;2;3: …3:42† The quantities ijˆ^e0 i^ejare called the coecients of transformation. They are the direction cosines of the primed coordinate axes relative to the unprimed ones ijˆ^e0 i^ejˆcos…x0 i;xj†; i;jˆ1;2;3: …3:42a† Eq. (3.42) can be written conveniently in the following matrix form x0 1 x0 2 x0 30 B@1 CAˆ111213 212223 3132330 B@1 CAx1 x2 x30 B@1 CA …3:43a† or ~X0ˆ~…†~X; …3:43b† where ~X0and ~Xare the column matrices, ~…†is called a transformation (or rotation) matrix; it acts as a linear operator which transforms the vector /C88into the vector /C880. Strictly speaking, we should describe the matrix ~…†as the matrix representation of the linear operator ^. The concept of linear operator is more general than that of matrix. Not all of the nine quantities ijare independent; six relations exist among the ij, hence only three of them are independent. These six relations are found by using the fact that the magnitude of the vector must be the same in both systems: X3 iˆ1…x0 i†2ˆX3 iˆ1x2 i: …3:44† With the help of Eq. (3.42), the left hand side of the last equation becomes X3 iˆ1X3 jˆ1ijxj/C32!X3 kˆ1ikxk/C32! ˆX3 iˆ1X3 jˆ1X3 kˆ1ijikxjxk; which, by rearranging the summations, can be rewritten as X3 kˆ1X3 jˆ1X3 iˆ1ijik/C32! xjxk: This last expression will reduce to the right hand side of Eq. (3.43) if and only if X3 iˆ1ijikˆjk; j;kˆ1;2;3: …3:45† 118MATRI/C88 ALGEBRA Eq. (3.45) gives six relations among the ij, and is known as the orthogonal condition. If the primed coordinates system is generated by a rotation about the x3-axis through an angle as shown in Fig. 3.2. Then from Example 3.5, we have x0 1ˆx1cos‡x2sin;x0 2ˆÿx1sin‡x2cos;x0 3ˆx3: …3:46† Thus 11ˆcos;  12ˆsin;  13ˆ0; 21ˆÿsin;  22ˆcos;  23ˆ0; 31ˆ0; 32ˆ0; 33ˆ1: We can also obtain these elements from Eq. (3.42a). It is obvious that only three of them are independent, and it is easy to check that they satisfy the condition given in Eq. (3.45). Now the rotation matrix takes the simple form ~…†ˆcossin0 ÿsincos0 00 10 B@1 CA …3:47† and its transpose is ~T…†ˆcosÿsin0 sin cos0 00 10 B@1 CA: Now take the product ~T…†~…†ˆcossin0 ÿsincos0 00 10 B@1 CAcosÿsin0 sin cos0 00 10 B@1 CAˆ100 0100010 B@1 CAˆ~I; which shows that the rotation matrix is an orthogonal matrix. In fact, rotation matrices are orthogonal matrices, not limited to ~…†of Eq. (3.47). The proof of this is easy. Since coordinate transformations are reversible by interchanging oldand new indices, we must have ~ ÿ1ÿ ijˆ^eold i^enewjˆ^enewj^eoldiˆjiˆ ~Tÿ ij: Hence rotation matrices are orthogonal matrices. It is obvious that the inverse of an orthogonal matrix is equal to its transpose. A rotation matrix such as given in Eq. (3.47) is a continuous function of its argument . So its determinant is also a continuous function of and, in fact, it is equal to 1 for any . There are matrices of coordinate changes with a determinant ofÿ1. These correspond to inversion of the coordinate axes about the origin and 119ROTATION MATRICES change the handedness of the coordinate system. Examples of such parity trans- formations are ~P1ˆÿ100 010 0010 B@1 CA; ~P3ˆÿ100 0ÿ10 00 ÿ10 B@1 CA; ~P2 iˆI: They change the signs of an odd number of coordinates of a fixed point rin space (Fig. 3.3). What is the advantage of using matrices in describing rotation in space/C63 One of the advantages is that successive transformations 1 ;2;...;mof the coordinate axes about the origin are described by successive matrix multiplications as far as their e/C128ects on the coordinates of a fixed point are concerned: If~X…1†ˆ~1~X;~X…2†ˆ~2~X…1†;...;then ~X…m†ˆ~m~X…mÿ1†ˆ… ~m~mÿ1 ~1†~Xˆ~R~X where ~Rˆ~m~mÿ1 ~1 is the resultant (or net) rotation matrix for the msuccessive transformations taken place in the specified manner. Example 3.9 Consider a rotation of the x1-,x2-axes about the x3-axis by an angle . If this rotation is followed by a back-rotation of the same angle in the opposite direction, 120MATRI/C88 ALGEBRA Figure 3.3. Parity transformations of the coordinate system. that is, by ÿ, we recover the original coordinate system. Thus ~R…ÿ†~R…†ˆ100 010 0010 B@1 CAˆ~Rÿ1…†~R…†: Hence ~Rÿ1…†ˆ ~R…ÿ†ˆcosÿsin0 sin cos0 00 10 B@1 CAˆ~RT…†; which shows that a rotation matrix is an orthogonal matrix. We would like to make one remark on rotation in space. In the above discus- sion, we have considered the vector to be fixed and rotated the coordinate axes. The rotation matrix can be thought of as an operator that, acting on the unprimed system, transforms it into the primed system. This view is often called the passive view of rotation. We could equally well keep the coordinate axes fixed and rotatethe vector through an equal angle, but in the opposite direction. Then the rotation matrix would be thought of as an operator acting on the vector, say /C88, and changing it into /C88 0. This procedure is called the active view of the rotation. /C84race of a matri/C120 Recall that the trace of a square matrix ~Ais defined as the sum of all the principal diagonal elements: Tr ~AˆX kakk: It can be proved that the trace of the product of a finite number of matrices isinvariant under any cyclic permutation of the matrices. We leave this as home work. /C79rthogonal and unitar/C121 transformations Eq. (3.42) is a linear transformation and it is called an orthogonal transformation, because the rotation matrix is an orthogonal matrix. One of the properties of anorthogonal transformation is that it preserves the length of a vector. A more useful linear transformation in physics is the unitary transformation: ~Yˆ~U~X …3:48† in which ~Xand ~Yare column matrices (vectors) of order n1 and ~Uis a unitary matrix of order nn. One of the properties of a unitary transformation is that it 121TRACE OF A MATRI/C88 preserves the norm of a vector. To see this, premultiplying Eq. (3.48) by ~Y/C121…ˆ ~X/C121~U/C121) and using the condition ~U/C121~Uˆ~I, we obtain ~Y/C121~Yˆ~X/C121~U/C121~U~Xˆ~X/C121~X …3:49a† or Xn kˆ1yk/C42ykˆXn kˆ1xk/C42xk: …3:49b† This shows that the norm of a vector remains invariant under a unitary transfor- mation. If the matrix ~Uof transformation happens to be real, then ~Uis also an orthogonal matrix and the transformation (3.48) is an orthogonal transformation,and Eqs. (3.49) reduce to ~Y T~Yˆ~XT~X; …3:50a† Xn kˆ1y2 kˆXn kˆ1x2k; …3:50b† as we expected. /C83imilarit/C121 transformation We now consider a di/C128erent linear transformation, the similarity transformation that, we shall see later, is very useful in diagonalization of a matrix. To get the idea about similarity transformations, we consider vectors rand/C82in a particular basis, the coordinate system Ox1x2x3, which are connected by a square matrix ~A: /C82ˆ~Ar: …3:51a† Now rotating the coordinate system about the origin Owe obtain a new system Ox0 1x0 2x0 3(a new basis). The vectors rand /C82have not been a/C128ected by this rotation. Their components, however, will have di/C128erent values in the new system,and we now have /C82 0ˆ~A0r0: …3:51b† The matrix ~A0in the new (primed) system is called similar to the matrix ~Ain the old (unprimed) system, since they perform same function. Then what is the rela-tionship between matrices ~Aand ~A 0/C63 This information is given in the form of coordinate transformation. We learned in the previous section that the com-ponents of a vector in the primed and unprimed systems are connected by a matrix equation similar to Eq. (3.43). Thus we have rˆ~Sr 0and /C82ˆ~S/C820; 122MATRI/C88 ALGEBRA where ~Sis a non-singular matrix, the transition matrix from the new coordinate system to the old system. With these, Eq. (3.51a) becomes ~S/C820ˆ~A~Sr0 or /C820ˆ~Sÿ1~A~Sr0: Combining this with Eq. (3.51) gives ~A0ˆ~Sÿ1~A~S; …3:52† where ~A0and ~Aare similar matrices. Eq. (3.52) is called a similarity transforma- tion. Generalization of this idea to n-dimensional vectors is straightforward. In this case, we take rand/C82as two n-dimensional vectors in a particular basis, having their coordinates connected by the matrix ~A(annsquare matrix) through Eq. (3.51a). In another basis they are connected by Eq. (3.51b). The relationship between ~Aand ~A0is given by Eq. (3.52). The transformation of ~Ainto ~Sÿ1~A~S is called a similarity transformation. All identities involving vectors and matrices will remain invariant under a similarity transformation since this arises only in connection with a change inbasis. That this is so can be seen in the following two simple examples. Example 3.10 Given the matrix equation ~A~Bˆ~C, and the matrices ~A,~B,~Csubjected to the same similarity transformation, show that the matrix equation is invariant. Solution: Since the three matrices are all subjected to the same similarity trans- formation, we have ~A 0ˆ~S~A~Sÿ1; ~B0ˆ~S~B~Sÿ1; ~C0ˆ~S~C~Sÿ1 and it follows that ~A0~B0ˆ… ~S~A~Sÿ1†…~S~B~Sÿ1†ˆ ~S~A~I~B~Sÿ1ˆ~S~A~B~Sÿ1ˆ~S~C~Sÿ1ˆ~C0: Example 3.11 Show that the relation ~A/C82ˆ~Bris invariant under a similarity transformation. Solution: Since matrices ~Aand ~Bare subjected to the same similarity transfor- mation, we have ~A0ˆ~S~A~Sÿ1; ~B0ˆ~S~B~Sÿ1 we also have /C820ˆ~S/C82; r0ˆ~Sr: 123SIMILARITY TRANSFORMATION Then ~A0/C820ˆ… ~S~A~Sÿ1†…S/C82†ˆ ~S~A/C82and ~B0r0ˆ… ~S~B~Sÿ1†…Sr†ˆ ~S~Br thus ~A0/C820ˆ~B0r0: We shall see in the following section that similarity transformations are very useful in diagonalization of a matrix, and that two similar matrices have the same eigenvalues. /C84he matri/C120 eigen/C118alue problem As we saw in preceding sections, a linear transformation generally carries a vector /C88ˆ…x1;x2;...;xn†into a vector /C89ˆ…y1;y2;...;yn†:However, there may exist certain non-zero vectors for which ~A/C88is just /C88multiplied by a constant  ~A/C88ˆ/C88: …3:53† That is, the transformation represented by the matrix (operator) ~Ajust multiplies the vector /C88by a number . Such a vector is called an eigenvector of the matrix ~A, and is called an eigenvalue (German: eigenwert ) or characteristic value of the matrix ~A. The eigenvector is said to ‘belong’ (or correspond) to the eigenvalue. And the set of the eigenvalues of a matrix (an operator) is called its eigenvaluespectrum. The problem of finding the eigenvalues and eigenvectors of a matrix is called an eigenvalue problem. We encounter problems of this type in all branches of physics, classical or quantum. Various methods for the approximate determina- tion of eigenvalues have been developed, but here we only discuss the fundamental ideas and concepts that are important for the topics discussed inthis book. There are two parts to every eigenvalue problem. First, we compute the eigen- value , given the matrix ~A. Then, we compute an eigenvector Xfor each previously computed eigenvalue . /C68etermination of eigenvalues and eigenvectors We shall now demonstrate that any square matrix of order nhas at least 1 and at most ndistinct (real or complex) eigenvalues. To this purpose, let us rewrite the system of Eq. (3.53) as …~Aÿ~I†Xˆ0: …3:54† This matrix equation really consists of nhomogeneous linear equations in the n unknown elements x iofX: 124MATRI/C88 ALGEBRA a11ÿ …† x1‡a12x2‡  ‡ a1nxnˆ0 a21x1‡a22ÿ …† x2‡  ‡ a2nxnˆ0 ... an1x1‡an2x2‡  ‡ annÿ …† xnˆ09 >>>>>= >>>>>;…3:55† In order to have a non-zero solution, we recall that the determinant of the coe- cients must be zero; that is, det…~Aÿ~I†ˆa 11ÿ a12  a1n a21 a22ÿ a2n ......... an1 an2 annÿ/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ0: …3:56† The expansion of the determinant gives an nth order polynomial equation in , and we write this as c 0n‡c1nÿ1‡c2nÿ2‡  ‡ cnÿ1‡cnˆ0; …3:57† where the coecients ciare functions of the elements ajkof ~A. Eq. (3.56) or (3.57) is called the characteristic equation corresponding to the matrix ~A. We have thus obtained a very important result: the eigenvalues of a square matrix ~Aare the roots of the corresponding characteristic equation (3.56) or (3.57). Some of the coecients cican be readily determined; by an inspection of Eq. (3.56) we find c0ˆ… ÿ 1†n;c1ˆ… ÿ 1†nÿ1…a11‡a22‡‡ ann†;cnˆdet ~A: …3:58† Now let us rewrite the characteristic polynomial in terms of its nroots 1;2;...;n c0n‡c1nÿ1‡c2nÿ2‡  ‡ cnÿ1‡cnˆ1ÿ …† 2ÿ … † nÿ …† ; then we see that c1ˆ… ÿ 1†nÿ1…1‡2‡‡ n†;cnˆ12n: …3:59† Comparing this with Eq. (3.58), we obtain the following two important results on the eigenvalues of a matrix: (1) The sum of the eigenvalues equals the trace (spur) of the matrix: 1‡2‡‡ nˆa11‡a22‡‡ annTr ~A: …3:60† (2) The product of the eigenvalues equals the determinant of the matrix: 12nˆdet ~A: …3:61† 125THE MATRI/C88 EIGENVALUE PROBLEM Once the eigenvalues have been found, corresponding eigenvectors can be found from the system (3.55). Since the system is homogeneous, if Xis an eigenvector of ~A, then kX, where kis any constant (not zero), is also an eigen- vector of ~Acorresponding to the same eigenvalue. It is very easy to show this. Since ~AXˆX, multiplying by an arbitrary constant kwill give k~AXˆkX. Now k~Aˆ~Ak (every matrix commutes with a scalar), so we have ~A…kX†ˆ…kX†;showing that kXis also an eigenvector of ~Awith the same eigenvalue . But kXis linearly dependent on X, and if we were to count all such eigenvectors separately, we would have an infinite number of them. Such eigenvectors are therefore not counted separately. A matrix of order ndoes not necessarily have nlinearly independent eigenvectors; some of them may be repeated. (This will happen when the char-acteristic polynomial has two or more identical roots.) If an eigenvalue occurs m times, mis called the multiplicity of the eigenvalue. The matrix has at most m linearly independent eigenvectors all corresponding to the same eigenvalue. Such linearly independent eigenvectors having the same eigenvalue are said to be degen- erate eigenvectors; in this case, m-fold degenerate. We will deal only with those matrices that have nlinearly independent eigenvectors and they are diagonalizable matrices. Example 3.12 Find ( a) the eigenvalues and ( b) the eigenvectors of the matrix ~Aˆ54 12 : Solution: (a) The eigenvalues: The characteristic equation is det…~Aÿ~I†ˆ5ÿ 4 12 ÿ/C12/C12/C12/C12/C12/C12/C12/C12ˆ 2ÿ7‡6ˆ0 which has two roots 1ˆ6 and 2ˆ1: (b) The eigenvectors: For ˆ1the system (3.55) assumes the form ÿx1‡4x2ˆ0; x1ÿ4x2ˆ0: Thus x1ˆ4x2, and X1ˆ4 1 126MATRI/C88 ALGEBRA is an eigenvector of ~Acorresponding to 1ˆ6. In the same way we find the eigenvector corresponding to 2ˆ1: X2ˆ1 ÿ1 : Example 3.13 If~Ais a non-singular matrix, show that the eigenvalues of ~Aÿ1are the reciprocals of those of ~Aand every eigenvector of ~Ais also an eigenvector of ~Aÿ1. Solution: Letbe an eigenvalue of ~Acorresponding to the eigenvector X,s o that ~AXˆX: Since ~Aÿ1exists, multiply the above equation from the left by ~Aÿ1 ~Aÿ1~AXˆ~Aÿ1X/C41Xˆ~Aÿ1X: Since ~Ais non-singular, must be non-zero. Now dividing the above equation by , we have ~Aÿ1Xˆ…1=†X: Since this is true for every value of ~A, the results follows. Example 3.14Show that all the eigenvalues of a unitary matrix have unit magnitude. Solution: Let ~Ube a unitary matrix and Xan eigenvector of ~Uwith the eigen- value , so that ~UXˆX: Taking the hermitian conjugate of both sides, we have X /C121~U/C121ˆ/C42X/C121: Multiplying the first equation from the left by the second equation, we obtain X/C121~U/C121~UXˆ/C42X/C121X: Since ~Uis unitary, ~U/C121~U/C61 ~I, so that the last equation reduces to X/C121X…jj2ÿ1†ˆ0: Now X/C121Xis the square of the norm of Xand hence cannot vanish unless Xis a null vector and so we must have jj2ˆ1o rjjˆ1;proving the desired result. 127THE MATRI/C88 EIGENVALUE PROBLEM Example 3.15 Show that similar matrices have the same characteristic polynomial and hence the same eigenvalues. (Another way of stating this is to say that the eigenvalues of a matrix are invariant under similarity transformations.) Solution: Let ~Aand ~Bbe similar matrices. Thus there exists a third matrix ~S such that ~Bˆ~Sÿ1~A~S. Substituting this into the characteristic polynomial of matrix ~Bwhich is j~Bÿ~Ij, we obtain j~BÿIjˆj ~Sÿ1~A~Sÿ~Ijˆj ~Sÿ1…~Aÿ~I†~Sj: Using the properties of determinants, we have j~Sÿ1…~Aÿ~I†~Sjˆj ~Sÿ1jj~Aÿ~Ijj~Sj: Then it follows that j~Bÿ~Ijˆj ~Sÿ1…~Aÿ~I†~Sjˆj ~Sÿ1jj~Aÿ~Ijj~Sjˆj ~Aÿ~Ij; which shows that the characteristic polynomials of ~Aand ~Bare the same; their eigenvalues will also be identical. /C69igen/C118alues and eigen/C118ectors of hermitian matrices In quantum mechanics complex variables are unavoidable because of the form of the Schro /C200dinger equation. And all quantum observables are represented by her- mitian operators. So physicists are almost always dealing with adjoint matrices,hermitian matrices, and unitary matrices. Why are physicists interested in hermi- tian matrices/C63 Because they have the following properties: (1) the eigenvalues of a hermitian matrix are real, and (2) its eigenvectors corresponding to distinct eigen-values are orthogonal, so they can be used as basis vectors. We now proceed to prove these important properties. (1) the eigenvalues of a hermitian matrix are real. Let ~Hbe a hermitian matrix and Xa non-trivial eigenvector corresponding to the eigenvalue , so that ~HXˆX: …3:62† Taking the hermitian conjugate and note that ~H /C121ˆ ~H, we have X/C121~Hˆ/C42X/C121: …3:63† Multiplying (3.62) from the left by X/C121, and (3.63) from the right by X/C121, and then subtracting, we get …ÿ/C42†X/C121Xˆ0: …3:64† Now, since X/C121Xcannot be zero, it follows that ˆ/C42, or that is real. 128MATRI/C88 ALGEBRA (2) The eigenvectors corresponding to distinct eigenvalues are orthogonal. LetX1andX2be eigenvectors of ~Hcorresponding to the distinct eigenvalues 1 and 2, respectively, so that ~HX 1ˆ1X1; …3:65† ~HX 2ˆ2X2: …3:66† Taking the hermitian conjugate of (3.66) and noting that /C42ˆ, we have X/C121 2~Hˆ2X/C121 2: …3:67† Multiplying (3.65) from the left by X/C121 2and (3.67) from the right by X1, then subtracting, we obtain …1ÿ2†X/C121 2‡X1ˆ0: …3:68† Since 1ˆ2, it follows that X/C121 2X1ˆ0 or that X1andX2are orthogonal. IfXis an eigenvector of ~H, any multiple of X,X, is also an eigenvector of ~H. Thus we can normalize the eigenvector Xwith a properly chosen scalar . This means that the eigenvectors of ~Hcorresponding to distinct eigenvalues are ortho- normal. Just as the three orthogonal unit coordinate vectors ^e1;^e2;and ^e3form the basis of a three-dimensional vector space, the orthonormal eigenvectors of ~H may serve as a basis for a function space. /C68iagonali/C122ation of a matri/C120 Let ~Aˆ…aij†be a square matrix of order n, which has nlinearly independent eigenvectors Xiwith the corresponding eigenvalues i:~AXiˆiXi. If we denote the eigenvectors Xiby column vectors with elements x1i;x2i;...;xni, then the eigenvalue equation can be written in matrix form: a11a12a1n a21a22a2n ......... an1an2ann0 BBBBB@1 CCCCCAx 1i x2i ... xni0 BBBBB@1 CCCCCAˆ ix1i x2i ... xni0 BBBBB@1 CCCCCA: …3:69† From the above matrix equation we obtain X n kˆ1ajkxkiˆixji: …3:69b† Now we want to diagonalize ~A. To this purpose, we can follow these steps. We first form a matrix ~Sof order nnwhose columns are the vector Xi, that is, 129DIAGONALI/C90ATION OF A MATRI/C88 ~Sˆx11x1ix1n x21x2ix2n ......... xn1xnixnn0 BBBBB@1 CCCCCA;…~S† ijˆxij: …3:70† Since the vectors Xiare linear independent, ~Sis non-singular and ~Sÿ1exists. We then form a matrix ~Sÿ1~A~S; this is a diagonal matrix whose diagonal elements are the eigenvalues of ~A. To show this, we first define a diagonal matrix ~Bwhose diagonal elements are i …iˆ1;2;...;n†: ~Bˆ1 2 ... n0 BBBBB@1 CCCCCA; …3:71† and we then demonstrate that ~S ÿ1~A~Sˆ~B: …3:72a† Eq. (3.72a) can be rewritten by multiplying it from the left by ~Sas ~A~Sˆ~S~B: …3:72b† Consider the left hand side first. Taking the jith element, we obtain …~A~S†jiˆXn kˆ1…~A†jk…~S†kiˆXn kˆ1ajkxki: …3:73a† Similarly, the jith element of the right hand side is …~S~B†jiˆXn kˆ1…~S†jk…~B†kiˆXn kˆ1xjkikiˆixji: …3:73b† Eqs. (3.73a) and (3.73b) clearly show the validity of Eq. (3.72a). It is important to note that the matrix ~Sthat is able to diagonalize matrix ~Ais not unique. This is because we could arrange the eigenvectors X1;X2;...;Xnin any order to construct ~S. We summarize the procedure for diagonalizing a diagonalizable nnmatrix ~A: Step 1. Find nlinearly independent eigenvectors of ~A;X1;X2;...;Xn. Step 2. Form the matrix ~Shaving X1;X2;...;Xnas its column vectors. Step 3. Find the inverse of ~S,~Sÿ1. Step 4. The matrix ~Sÿ1~A~Swill then be diagonal with 1;2;...;nas its succes- sive diagonal elements, where iis the eigenvalue corresponding to Xi. 130MATRI/C88 ALGEBRA Example 3.16 Find a matrix ~Sthat diagonalizes ~Aˆ3ÿ20 ÿ23 0 00 50 B@1 CA: Solution: We have first to find the eigenvalues and the corresponding eigen- vectors of matrix ~A. The characteristic equation of ~Ais 3ÿÿ20 ÿ23 ÿ 0 00 5 ÿ/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ…ÿ1†…ÿ5† 2ˆ0; so that the eigenvalues of ~Aareˆ1 and ˆ5. By definition ~Xˆx1 x2 x30 B@1 CA is an eigenvector of ~Acorresponding to if and only if ~Xis a non-trivial solution of (~Iÿ~A†~Xˆ0, that is, of ÿ32 0 2 ÿ30 00 ÿ50 B@1 CAx1 x2 x30 B@1 CAˆ0 000 B@1 CA: Ifˆ5 the above equation becomes 220 2200000 B@1 CAx 1 x2 x30 B@1 CAˆ0 000 B@1 CAor2x 1‡2x2‡0x3 2x1‡2x2‡0x3 0x1‡0x2‡0x30 B@1 CAˆ0 000 B@1 CA: Solving this system yields x 1ˆÿs;x2ˆs;x3ˆt; where sand tare arbitrary values. Thus the eigenvectors of ~Acorresponding to ˆ5 are the non-zero vectors of the form ~Xˆÿs s t0 B@1 CAˆÿs s 00 B@1 CA‡0 0 t0 B@1 CAˆsÿ1 1 00 B@1 CA‡t0 0 10 B@1 CA: 131DIAGONALI/C90ATION OF A MATRI/C88 Since ÿ1 1 00 B@1 CAand0 0 10 B@1 CA are linearly independent, they are the eigenvectors corresponding to ˆ5. For ˆ1, we have ÿ22 0 2ÿ20 00 ÿ40 B@1 CAx1 x2 x30 B@1 CAˆ0 000 B@1 CAorÿ2x 1‡2x2‡0x3 2x1ÿ2x2‡0x3 0x1‡0x2ÿ4x30 B@1 CAˆ0 000 B@1 CA: Solving this system yields x 1ˆt;x2ˆt;x3ˆ0; where tis arbitrary. Thus the eigenvectors corresponding to ˆ1 are non-zero vectors of the form ~Xˆt t 00 B@1 CAˆt1 1 00 B@1 CA: It is easy to check that the three eigenvectors ~X1ˆÿ1 1 00 B@1 CA; ~X2ˆ0 0 10 B@1 CA; ~X3ˆ1 1 00 B@1 CA; are linearly independent. We now form the matrix ~Sthat has ~X1,~X2,a n d ~X3as its column vectors: ~Sˆÿ101 101 0100 B@1 CA: The matrix ~Sÿ1~A~Sis diagonal: ~Sÿ1~A~Sˆÿ1=21 =20 00 1 1=21 =200 B@1 CA3ÿ20 ÿ23 0 00 50 B@1 CAÿ101 101 0100 B@1 CAˆ500 050 0010 B@1 CA: There is no preferred order for the columns of ~S. If had we written ~Sˆÿ110 1100010 B@1 CA 132MATRI/C88 ALGEBRA then we would have obtained (verify) ~Sÿ1~A~Sˆ500 010 0010 B@1 CA: Example 3.17 Show that the matrix ~Aˆÿ32 ÿ21 is not diagonalizable. Solution: The characteristic equation of ~Ais ‡3ÿ2 2 ÿ1/C12/C12/C12/C12/C12/C12/C12/C12ˆ…‡1†2ˆ0: Thus ˆÿ1 the only eigenvalue of ~A; the eigenvectors corresponding to ˆÿ1 are the solutions of ‡3ÿ2 2 ÿ1x1 x2/C32! ˆ0 0 /C412ÿ2 2ÿ2x1 x2/C32! ˆ00 from which we have 2x 1ÿ2x2ˆ0; 2x1ÿ2x2ˆ0: The solutions to this system are x1ˆt;x2ˆt; hence the eigenvectors are of the form tt ˆt11 : Adoes not have two linearly independent eigenvectors, and is therefore not diagonalizable. /C69igen/C118ectors of commuting matrices There is a theorem on eigenvectors of commuting matrices that is of great impor- tance in matrix algebra as well as in quantum mechanics. This theorem states that: Two commuting matrices possess a common set of eigenvectors. 133EIGENVECTORS OF COMMUTING MATRICES We now proceed to prove it. Let ~Aand ~Bbe two square matrices, each of order n, which commute with each other, that is, ~A~Bÿ~B~Aˆ‰ ~A;~BŠˆ0: First, let be an eigenvalue of ~Awith multiplicity 1, corresponding to the eigen- vector X, so that ~AXˆX: …3:74† Multiplying both sides from the left by ~B ~B~AXˆ~BX: Because ~B~Aˆ~A~B, we have ~A…~BX†ˆ… ~BX†: Now ~Bis an nnmatrix and Xis an n1 vector; hence ~BXis also an n1 vector. The above equation shows that ~BXis also an eigenvector of ~Awith the eigenvalue . Now Xis a non-degenerate eigenvector of ~A, any other vector which is an eigenvector of ~Awith the same eigenvalue as that of Xmust be multiple of X. Accordingly ~BXˆ/C22X; where /C22is a scalar. Thus we have proved that: If two matrices commute, every non-degenerate eigenvector of one is also an eigenvector of the other, and vice versa. Next, let be an eigenvalue of ~Awith multiplicity k.S o ~Ahasklinearly inde- pendent eigenvectors, say X1;X2;...;Xk, each corresponding to : ~AXiˆXi;1ik: Multiplying both sides from the left by ~B, we obtain ~A…~BXi†ˆ…~BXi†; which shows again that ~BXis also an eigenvector of ~Awith the same eigenvalue . /C67a/C121le/C121/C177/C72amilton theorem The Cayley–Hamilton theorem is useful in evaluating the inverse of a square matrix. We now introduce it here. As given by Eq. (3.57), the characteristic equation associated with a square matrix ~Aof order nmay be written as a poly- nomial f…†ˆXn iˆ0cinÿiˆ0; 134MATRI/C88 ALGEBRA where are the eigenvalues given by the characteristic determinant (3.56). If we replace inf…†by the matrix ~Aso that f…~A†ˆXn iˆ0ci~Anÿi: The Cayley–Hamilton theorem says that f…~A†ˆ0o rXn iˆ0ci~Anÿiˆ0; …3:75† that is, the matrix ~Asatisfies its characteristic equation. We now formally multiply Eq. (3.75) by ~Aÿ1so that we obtain ~Aÿ1f…~A†ˆc0~Anÿ1‡c1~Anÿ2‡‡ cnÿ1~I‡cn~Aÿ1ˆ0: Solving for ~Aÿ1gives ~Aÿ1ˆÿ1 cnXnÿ1 iˆ0ci~Anÿ1ÿi"# ; …3:76† we can use this to find ~Aÿ1(Problem 3.28). Moment of inertia matri/C120 We shall see that physically diagonalization amounts to a simplification of the problem by a better choice of variable or coordinate system. As an illustrative example, we consider the moment of inertia matrix ~Iof a rotating rigid body (see Fig. 3.4). A rigid body can be considered to be a many-particle system, with the 135MOMENT OF INERTIA MATRI/C88 Figure 3.4. A rotating rigid body. distance between any particle pair constant at all times. Then its angular momen- tum about the origin Oof the coordinate system is LˆX m r /C118 ˆX m r …xr † where the subscript refers to mass malocated at r ˆ…x 1;x 2;x 3†, andxthe angular velocity of the rigid body. Expanding the vector triple product by using the vector identity /C65…/C66/C67†ˆ/C66…/C65/C67†ÿ/C67…/C65/C66†; we obtain LˆX m /C98r2 xÿr …r x†/C99: In terms of the components of the vectors r andx, the ith component of Liis LiˆX m /C33iX3 kˆ1x2 ;kÿx ;iX3 jˆ1x ;j/C33j"# ˆX j/C33jX m ijX kx2 ;kÿx ;ix ;j"# ˆX jIij/C33j or ~Lˆ~I~/C33: Both ~Land ~/C33are three-dimensional column vectors, while ~Iis a 3 3 matrix and is called the moment inertia matrix. In general, the angular momentum vector Lof a rigid body is not always parallel to its angular velocity xand ~Iis not a diagonal matrix. But we can orient the coordinate axes in space so that all the non-diagonal elements Iij…i6ˆj† vanish. Such special directions are called the principal axes of inertia. If the angular velocity is along one of these principal axes, the angular momentum and the angular velocity will be parallel. In many simple cases, especially when symmetry is present, the principal axes of inertia can be found by inspection. Normal modes of /C118ibrations Another good illustrative example of the application of matrix methods in classi-cal physics is the longitudinal vibrations of a classical model of a carbon dioxide molecule that has the chemical structure O–C–O. In particular, it provides a good example of the eigenvalues and eigenvectors of an asymmetric real matrix. 136MATRI/C88 ALGEBRA We can regard a carbon dioxide molecule as equivalent to a set of three par- ticles jointed by elastic springs (Fig. 3.5). Clearly the system will vibrate in some manner in response to an external force. For simplicity we shall consider only longitudinal vibrations, and the interactions of the oxygen molecules with oneanother will be neglected, so we consider only nearest neighbor interaction. The Lagrangian function Lfor the system is Lˆ 1 2m…_x2 1‡_x23†‡1 2M_x2 2ÿ1 2k…x2ÿx1†2ÿ12k…x3ÿx2†2; substituting this into Lagrange’s equations d dt/C64L /C64_xi ÿ/C64L /C64xiˆ0…iˆ1;2;3†; we find the equations of motion to be /C127x1ˆÿk m…x1ÿx2†ˆÿk mx1‡k mx2; /C127x2ˆÿk M…x2ÿx1†ÿk M…x2ÿx3†ˆk Mx1ÿ2k Mx2‡k Mx3; /C127x3ˆk mx2ÿk mx3; where the dots denote time derivatives. If we define ~Xˆx1 x2 x30 B@1 CA; ~Aˆÿk mk m0 ÿk Mÿ2k Mk M 0k mÿk m0 BBBBBBB@1 CCCCCCCA and, furthermore, if we define the derivative of a matrix to be the matrix obtained by di/C128erentiating each matrix element, then the above system of di/C128erential equa- tions can be written as /C127~Xˆ~A~X: 137NORMAL MODES OF VIBRATIONS Figure 3.5. A linear symmetrical carbon dioxide molecule. This matrix equation is reminiscent of the single di/C128erential equation /C127xˆax, with aa constant. The latter always has an exponential solution. This suggests that we try ~Xˆ~Ce/C33t; where /C33is to be determined and ~CˆC1 C2 C30 B@1 CA is an as yet unknown constant matrix. Substituting this into the above matrix equation, we obtain a matrix-eigenvalue equation ~A~Cˆ/C332~C or ÿk mk m0 ÿk Mÿ2k Mk M 0k mÿk m0 BBBBBBBBB@1 CCCCCCCCCAC 1 C2 C30 B@1 CAˆ/C332C1 C2 C30 B@1 CA: …3:77† Thus the possible values of /C33are the square roots of the eigenvalues of the asymmetric matrix ~Awith the corresponding solutions being the eigenvectors of the matrix ~A. The secular equation is ÿk mÿ/C332 k m0 ÿk Mÿ2k Mÿ/C332 k M 0k mÿk mÿ/C332/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ0: This leads to /C33 2ÿ/C332‡k m ÿ/C332‡k m‡2k M ˆ0: The eigenvalues are /C332ˆ0;k m;andk m‡2k M; 138MATRI/C88 ALGEBRA all real. The corresponding eigenvectors are determined by substituting the eigen- values back into Eq. (3.77) one eigenvalue at a time: (1) Setting /C332ˆ0 in Eq. (3.77) we find that C1ˆC2ˆC3. Thus this mode is not an oscillation at all, but is a pure translation of the system as a whole, no relative motion of the masses (Fig. 3.6( a)). (2) Setting /C332ˆk=min Eq. (3.77), we find C2ˆ0 and C3ˆÿC1. Thus the center mass Mis stationary while the outer masses vibrate in opposite directions with the same amplitude (Fig. 3.6( b)). (3) Setting /C332ˆk=m‡2k=Min Eq. (3.77), we find C1ˆC3, and C2ˆÿ2C1…m=M†. In this mode the two outer masses vibrate in unison and the center mass vibrates oppositely with di/C128erent amplitude (Fig. 3.6( c)). /C68irect product of matrices Sometimes the direct product of matrices is useful. Given an mmmatrix ~A and an nnmatrix ~B, the direct product of ~Aand ~Bis an mnmnmatrix, defined by ~Cˆ~A/C10~Bˆa11~Ba 12~B a1m~B a21~Ba 22~B a2m~B ......... am1~Ba m2~Bamm~B0 BBBBB@1 CCCCCA: For example, if ~Aˆa 11a12 a21a22/C32! ; ~Bˆb11b12 b21b22/C32! ; then 139DIRECT PRODUCT OF MATRICES Figure 3.6. Longitudinal vibrations of a carbon dioxide molecule. ~A/C10~Bˆa11~Ba 12~B a21~Ba 22~B/C32! ˆa11b11a11b12a12b11a12b12 a11b21a11b22a12b21a12b22 a21b11a21b12a22b11a22b12 a21b21a21b22a22b21a22b220 BBBB@1 CCCCA: Problems 3.1 For the pairs ~Aand ~Bgiven below, find ~A‡~B,~A~B,a n d ~A 2: ~Aˆ12 34 ; ~Bˆ5678 : 3.2 Show that an n-rowed diagonal matrix ~D ~Dˆk0 0 0k 0 ......... k0 BBBB@1 CCCCA commutes with any n-rowed square matrix ~A:~A~Dˆ~D~Aˆk~A. 3.3 If ~A,~B,a n d ~Care any matrices such that the addition ~B‡~Cand the products ~A~Band ~A~Care defined, show that ~A(~B‡~C†ˆ ~A~B/C43~A~C. That is, that matrix multiplication is distributive. 3.4 Given ~Aˆ010 101 0100 B@1 CA; ~Bˆ100 010 0010 B@1 CA; ~Cˆ10 0 00 0 00 ÿ10 B@1 CA; show that /C91 ~A;~BŠˆ0, and /C91 ~B;~CŠˆ0, but that ~Adoes not commute with ~C. 3.5 Prove that ( ~A‡~B† Tˆ~AT‡~BT. 3.6 Given ~Aˆ2ÿ3 04 ; ~Bˆÿ52 21 ;and ~Cˆ01 ÿ2 30 4 /C58 (a) Find 2 ~Aÿ4~B,2 ( ~Aÿ2~B) (b) Find ~AT;~BT;…~BT†T (c) Find ~CT;…~CT†T (d)I s ~A‡~Cdefined/C63 (e)I s ~C‡~CTdefined/C63 (f)I s ~A‡~ATsymmetric/C63 140MATRI/C88 ALGEBRA (g)I s ~Aÿ~ATantisymmetric/C63 3.7 Show that the matrix ~Aˆ140 250 3600 B@1 CA is not invertible. 3.8 Show that if ~Aand ~Bare invertible matrices of the same order, then ~A~Bis invertible. 3.9 Given ~Aˆ123 2531080 B@1 CA; find ~A ÿ1and check the answer by direct multiplication. 3.10 Prove that if ~Ais a non-singular matrix, then det( ~Aÿ1†ˆ1=det…~A). 3.11 If ~Ais an invertible nnmatrix, show that ~AXˆ0 has only the trivial solution. 3.12 Show, by computing a matrix inverse, that the solution to the following system is x1ˆ4,x2ˆ1: x1ÿx2ˆ3; x1‡x2ˆ5: 3.13 Solve the system ~AXˆ~Bif ~Aˆ100 020 0010 B@1 CA; ~Bˆ1 2 30 B@1 CA: 3.14 Given matrix ~A, find A/C42, AT, and A/C121, where ~Aˆ2‡3i1ÿi 5i ÿ3 1‡i6ÿi1‡3iÿ1ÿ2i 5ÿ6i30 ÿ40 B@1 CA: 3.15 Show that: (a) The matrix ~A~A/C121, where ~Ais any matrix, is hermitian. (b)…~A~B†/C121ˆ~B/C121~A/C121: (c)I f ~A;~Bare hermitian, then ~A~B‡~B~Ais hermitian. (d)I f ~Aand ~Bare hermitian, then i…~A~Bÿ~B~A†is hermitian. 3.16 Obtain the most general orthogonal matrix of order 2. /C91Hint: use relations (3.34a) and (3.34b)./C93 3.17. Obtain the most general unitary matrix of order 2. 141PROBLEMS 3.18 If ~A~Bˆ0, show that one of these matrices must have zero determinant. 3.19 Given the Pauli spin matrices (which are very important in quantum mechanics) 1ˆ01 10 ; 2ˆ0ÿi i0 ; 3ˆ100ÿ1 ; (note that the subscripts x;y, and zare sometimes used instead of 1, 2, and 3). Show that (a) they are hermitian, (b) 2 iˆ~I;iˆ1;2;3 (c) as a result of ( a) and ( b) they are also unitary, and (d)/C911;2Šˆ2I3et cycl . Find the inverses of 1;2;3: 3.20 Use a rotation matrix to show that sin…1‡2†ˆsin1cos2‡sin2cos1: 3.21 Show that: Tr ~A~BˆTr ~B~Aand Tr ~A~B~CˆTr ~B~C~AˆTr ~C~A~B: 3.22 Show that: ( a) the trace and ( b) the commutation relation between two matrices are invariant under similarity transformations. 3.23 Determine the eigenvalues and eigenvectors of the matrix ~Aˆab ÿba : Given ~Aˆ57 ÿ5 04 ÿ1 28 ÿ30 B@1 CA; find a matrix ~Sthat diagonalizes ~A, and show that ~Sÿ1~A~Sis diagonal. 3.25 If ~Aand ~Bare square matrices of the same order, then det( ~A~B†ˆdet…~A†det…~B†:Verify this theorem if ~Aˆ2ÿ1 32 ; ~Bˆ72 ÿ34 : 3.26 Find a common set of eigenvectors for the two matrices ~Aˆÿ1 6p  2p 6p 03p 2p  3p ÿ20 B@1 CA; ~Bˆ10 6p ÿ 2p 6p 93p ÿ2p  3p 110 B@1 CA: 3.27 Show that two hermitian matrices can be made diagonal if and only if they commute. 142MATRI/C88 ALGEBRA 3.28 Show the validity of the Cayley–Hamilton theorem by applying it to the matrix ~Aˆ54 12 ; then use the Cayley–Hamilton theorem to find the inverse of the matrix ~A. 3.29 Given ~Aˆ01 10 ; ~Bˆ0ÿi i0 ; find the direct product of these matrices, and show that it does not com- mute. 143PROBLEMS 4 /C70ourier series and integrals Fourier series are infinite series of sines and cosines which are capable of repre- senting almost any periodic function whether continuous or not. Periodic func- tions that occur in physics and engineering problems are often very complicated and it is desirable to represent them in terms of simple periodic functions. Therefore the study of Fourier series is a matter of great practical importance for physicists and engineers. The first part of this chapter deals with Fourier series. Basic concepts, facts, and techniques in connection with Fourier series will be introduced and developed, along with illustrative examples. They are followed by Fourier integrals and Fourier transforms. Periodic functions If function f…x†is defined for all xand there is some positive constant Psuch that f…x‡P†ˆf…x†… 4:1† then we say that f…x†is periodic with a period P(Fig. 4.1). From Eq. (4.1) we also 144Figure 4.1. A general periodic function. have, for all xand any integer n, f…x‡nP†ˆf…x†: That is, every periodic function has arbitrarily large periods and contains arbi- trarily large numbers in its domain. We call Pthe fundamental (or least) period, or simply the period. A periodic function need not be defined for all values of its independent vari- able. For example, tan xis undefined for the values xˆ…=2†‡n. But tan xis a periodic function in its domain of definition, with as its fundamental period: tan(x‡†ˆtanx. Example 4.1 (a) The period of sin xis 2, since sin( x‡2†, sin(x‡4†;sin…x‡6†;...are all equal to sin x, but 2 is the least value of P. And, as shown in Fig. 4.2, the period of sin nxis 2=n, where nis a positive integer. (b) A constant function has any positive number as a period. Since f…x†ˆ c(const.) is defined for all real x, then, for every positive number P, f…x‡P†ˆcˆf…x†. Hence Pis a period of f. Furthermore, fhas no fundamental period. (c) f…x†ˆ/C75for 2 nx…2n‡1† ÿ/C75for…2n‡1†x<…2n‡2†( nˆ0;1;2;3;... is periodic of period 2 (Fig. 4.3). 145PERIODIC FUNCTIONS Figure 4.2. Sine functions. Figure 4.3. A square wave function. Fourier series/C59 /C69uler/C177Fourier formulas If the general periodic function f…x†is defined in an interval ÿx, the Fourier series of f…x†in /C91ÿ; /C93 is defined to be a trigonometric series of the form f…x†ˆ1 2a0‡a1cosx‡a2cos 2x‡‡ ancosnx‡ ‡b1sinx‡b2sin 2x‡‡ bnsinnx‡ ; …4:2† where the numbers a0;a1;a2;...;b1;b2;b3;...are called the Fourier coecients of f…x†in‰ÿ; Š. If this expansion is possible, then our power to solve physical problems is greatly increased, since the sine and cosine terms in the series can be handled individually without diculty. Joseph Fourier (1768–1830), a French mathematician, undertook the systematic study of such expansions. In 1807 he submitted a paper (on heat conduction) to the Academy of Sciences in Paris and claimed that every function defined on the closed interval ‰ÿ; Šcould be repre- sented in the form of a series given by Eq. (4.2); he also provided integral formulasfor the coecients a nandbn. These integral formulas had been obtained earlier by Clairaut in 1757 and by Euler in 1777. However, Fourier opened a new avenue byclaiming that these integral formulas are well defined even for very arbitrary functions and that the resulting coecients are identical for di/C128erent functions that are defined within the interval. Fourier’s paper was rejected by the Academy on the grounds that it lacked mathematical rigor, because he did not examine thequestion of the convergence of the series. The trigonometric series (4.2) is the only series which corresponds to f…x†. Questions concerning its convergence and, if it does, the conditions under which it converges to f…x†are many and dicult. These problems were partially answered by Peter Gustave Lejeune Dirichlet (German mathematician, 1805–1859) and will be discussed briefly later. Now let us assume that the series exists, converges, and may be integrated term by term. Multiplying both sides by cos mx, then integrating the result from ÿto ,w eh a v e Z  ÿf…x†cosmx dx ˆa0 2Z ÿcosmx dx ‡X1 nˆ1anZ ÿcosnxcosmx dx ‡X1 nˆ1bnZ ÿsinnxcosmx dx : …4:3† Now, using the following important properties of sines and cosines: Z ÿcosmx dx ˆZ ÿsinmx dx ˆ0i f mˆ1;2;3;...; Z ÿcosmxcosnx dx ˆZ ÿsinmxsinnx dx ˆ0i f n6ˆm; ifnˆm;( 146FOURIER SERIES AND INTEGRALS Z ÿsinmxcosnx dx ˆ0;for all m;n/C620; we find that all terms on the right hand side of Eq. (4.3) except one vanish: anˆ1 Z ÿf…x†cosnx dx ;nˆintegers ; …4:4a† the expression for a0can be obtained from the general expression for anby setting nˆ0. Similarly, if Eq. (4.2) is multiplied through by sin mxand the result is integrated fromÿto, all terms vanish save that involving the square of sin nx, and so we have bnˆ1 Z ÿf…x†sinnx dx : …4:4b† Eqs. (4.4a) and (4.4b) are known as the Euler–Fourier formulas. From the definition of a definite integral it follows that, if f…x†is single-valued and continuous within the interval ‰ÿ; Šor merely piecewise continuous (con- tinuous except at a finite numbers of finite jumps in the interval), the integrals in Eqs. (4.4) exist and we may compute the Fourier coecients of f…x†by Eqs. (4.4). If there exists a finite discontinuity in f…x†at the point x0(Fig. 4.1), the coe- cients a0;an;bnare determined by integrating first to xˆx0and then from x0to, as anˆ1 Zx0 ÿf…x†cosnx dx ‡Z x0f…x†cosnx dx ; …4:5a† bnˆ1 Zx0 ÿf…x†sinnx dx ‡Z x0f…x†sinnx dx : …4:5b† This procedure may be extended to any finite number of discontinuities. Example 4.2 Find the Fourier series which represents the function f…x†ˆÿkÿ<x<0 ‡k0<x<and f…x‡2†ˆf…x†;/C26 in the interval ÿx. 147FOURIER SERIES; EULER–FOURIER FORMULAS Solution: The Fourier coecients are readily calculated: anˆ1 Z0 ÿ…ÿk†cosnx dx ‡Z 0kcosnx dx ˆ1 ÿksinnx n/C12/C12/C12/C120 ÿ‡ksinnx n/C12/C12/C12/C12 0 ˆ0" bnˆ1 Z0 ÿ…ÿk†sinnx dx ‡Z 0ksinnx dx ˆ1 kcosnx n/C12/C12/C12/C120 ÿÿkcosnx n/C12/C12/C12/C12 0 ˆ2k n…1ÿcosn†" Now cos nˆÿ1 for odd n, and cos nˆ1 for even n. Thus b1ˆ4k=; b2ˆ0;b3ˆ4k=3;b4ˆ0;b5ˆ4k=5;... and the corresponding. Fourier series is 4k sinx‡1 3sin 3x‡15sin 5x‡  : For the special case kˆ=2, the Fourier series becomes 2 sinx‡23sin 3x‡25sin 5x‡ : The first two terms are shown in Fig. 4.4, the solid curve is their sum. We will see that as more and more terms in the Fourier series expansion are included, the sum more and more nearly approaches the shape of f…x†. This will be further demon- strated by next example. Example 4.3 Find the Fourier series that represents the function defined by f…t†ˆ0; ÿ<t<0 sint; 0<t</C26 in the interval ÿ<t< : 148FOURIER SERIES AND INTEGRALS Solution: anˆ1 Z0 ÿ0cosnt dt‡Z 0sintcosnt dt ˆÿ1 2cos…1ÿn†t 1ÿn‡cos…1‡n†t 1‡n  0ˆcosn‡1 …1ÿn†2; n6ˆ1/C12/C12/C12/C12/C12; a 1ˆ1 Z 0sintcostd tˆ1 sin2t 2/C12/C12/C12/C12 0ˆ0; bnˆ1 Z0 ÿ0sinnt dt‡Z 0sintsinnt dt ˆ1 2sin…1ÿn†t 1ÿnÿsin…1‡n†t 1‡n 0ˆ0 b1ˆ1 Z 0sin2td tˆ1 t 2ÿsin 2t 40 ˆ1 2: Accordingly the Fourier expansion of f…t†in /C91ÿ; /C93 may be written f…t†ˆ1 ‡sint 2ÿ2 cos 2t 3‡cos 4t 15‡cos 6t 35‡cos 8t 63‡  : The first three partial sums Sn…nˆ1;2;3) are shown in Fig. 4.5: S1ˆ1=; S2ˆ1=‡sint=2, and S3ˆ1=‡sin…t†=2ÿ2 cos…2t†=3: 149FOURIER SERIES; EULER–FOURIER FORMULAS Figure 4.4. The first two partial sums. /C71ibb/C39s phenomena From Figs. 4.4 and 4.5, two features of the Fourier expansion should be noted: (a) at the points of the discontinuity, the series yields the mean value; (b) in the region immediately adjacent to the points of discontinuity, the expan- sion overshoots the original function. This e/C128ect is known as the /C71ibb/C39s phenomena and occurs in all order of approximation. /C67on/C118ergence of Fourier series and /C68irichlet conditions The serious question of the convergence of Fourier series still remains: if we determine the Fourier coecients an;bnof a given function f…x†from Eq. (4.4) and form the Fourier series given on the right hand side of Eq. (4.2), will itconverge toward f…x†/C63 This question was partially answered by Dirichlet. Here is a restatement of the results of his study, which is often calledDirichlet’s theorem: (1) If f…x†is defined and single-valued except at a finite number of point in ‰ÿ; Š, (2) if f…x†is periodic outside ‰ÿ; Šwith period 2 (that is, f…x‡2†ˆf…x††, and (3) if f…x†andf 0…x†are piecewise continuous in ‰ÿ; Š, 150FOURIER SERIES AND INTEGRALS Figure 4.5. The first three partial sums of the series. then the series on the right hand side of Eq. (4.2), with coecients anandbngiven by Eqs. (4.4), converges to (i)f…x†,i fxis a point of continuity, or (ii)1 2‰f…x‡0†‡f…xÿ0†Š,i fxis a point of discontinuity as shown in Fig. 4.6, where f…x‡0†andf…xÿ0†are the right and left hand limits of f…x†atxand represent lim /C34!0f…x‡/C34†and lim /C34!0f…xÿ/C34†respectively, where /C34/C620. The proof of Dirichlet’s theorem is quite technical and is omitted in this treat- ment. The reader should remember that the Dirichlet conditions (1), (2), and (3) imposed on f…x†are sucient but not necessary. That is, if the above conditions are satisfied the convergence is guaranteed; but if they are not satisfied, the seriesmay or may not converge. The Dirichlet conditions are generally satisfied in practice. /C72alf-range Fourier series Unnecessary work in determining Fourier coecients of a function can beavoided if the function is odd or even. A function f…x†is called odd if f…ÿx†ˆÿ f…x†and even if f…x†f…ÿx†ˆf…x†. It is easy to show that in the Fourier series corresponding to an odd function f o…x†, only sine terms can be present in the series expansion in the interval ÿ<x<, for anˆ1 Z ÿfo…x†cosnx dx ˆ1 Z0 ÿfo…x†cosnx dx ‡Z 0fo…x†cosnx dx ˆ1 ÿZ 0fo…x†cosnx dx ‡Z 0fo…x†cosnx dx ˆ0 nˆ0;1;2;...;…4:6a† 151HALF-RANGE FOURIER SERIES Figure 4.6. A piecewise continuous function. but bnˆ1 Z0 ÿfo…x†sinnx dx ‡Z 0fo…x†sinnx dx ˆ2 Z 0fo…x†sinnx dx n ˆ1;2;3;...: …4:6b† Here we have made use of the fact that cos( ÿnx†ˆcosnxand sin …ÿnx†ˆ ÿsinnx. Accordingly, the Fourier series becomes fo…x†ˆb1sinx‡b2sin 2x‡ : Similarly, in the Fourier series corresponding to an even function fe…x†, only cosine terms (and possibly a constant) can be present. Because in this case, fe…x†sinnxis an odd function and accordingly bnˆ0 and the anare given by anˆ2 Z 0fe…x†cosnx dx n ˆ0;1;2;...: …4:7† Note that the Fourier coecients anandbn, Eqs. (4.6) and (4.7) are computed in the interval (0, ) which is halfof the interval ( ÿ; ). Thus, the Fourier sine or cosine series in this case is often called a half-range Fourier series. Any arbitrary function (neither even nor odd) can be expressed as a combina- tion of fe…x†andfo…x†as f…x†ˆ1 2f…x†‡f…ÿx† ‰Š ‡12f…x†ÿf…ÿx† ‰Š ˆ fe…x†‡fo…x†: When a half-range series corresponding to a given function is desired, the function is generally defined in the interval (0, ) and then the function is specified as odd or even, so that it is clearly defined in the other half of the interval …ÿ;0†. /C67hange of inter/C118al A Fourier expansion is not restricted to such intervals as ÿ<x< and 0<x<. In many problems the period of the function to be expanded may be some other interval, say 2 L. How then can the Fourier series developed above be applied to the representation of periodic functions of arbitrary period/C63 The problem is not a dicult one, for basically all that is involved is to change the variable. Let zˆ Lx …4:8a† then f…z†ˆf…x=L†ˆ/C70…x†: …4:8b† Thus, if f…z†is expanded in the interval ÿ<z<, the coecients being deter- mined by expressions of the form of Eqs. (4.4a) and (4.4b), the coecients for the 152FOURIER SERIES AND INTEGRALS expansion of /C70…x†in the interval ÿL<x<Lmay be obtained merely by sub- stituting Eqs. (4.8) into these expressions. We have then anˆ1 LZL ÿL/C70…x†cosn Lxd x n ˆ0;1;2;3;...; …4:9a† bnˆ1 LZL ÿL/C70…x†sinn Lxd x; nˆ1;2;3;...: …4:9b† The possibility of having expanding functions in which the period is other than 2increases the usefulness of Fourier expansion. As an example, consider the value of L, it is obvious that the larger the value of L, the larger the basic period of the function being expanded. As L!1 , the function would not be periodic at all. We will see later that in such cases the Fourier series becomes a Fourier integral. Parse/C118al/C39s identit/C121 Parseval’s identity states that: 1 2LZL ÿL‰f…x†Š2dxˆa0 22 ‡1 2X1 nˆ1…a2 n‡b2n†; …4:10† ifanandbnare coecients of the Fourier series of f…x†and if f…x†satisfies the Dirichlet conditions. It is easy to prove this identity. Assuming that the Fourier series corresponding tof…x†converges to f…x† f…x†ˆa0 2‡X1 nˆ1ancosnx L‡bnsinnx L : Multiplying by f…x†and integrating term by term from ÿLtoL, we obtain ZL ÿL‰f…x†Š2dxˆa0 2ZL ÿLf…x†dx ‡X1 nˆ1anZL ÿLf…x†cosnx Ldx‡bnZL ÿLf…x†sinnx Ldx/C26/C27 ˆa20 2L‡LX1 nˆ1a2 n‡b2nÿ ; …4:11† where we have used the results ZL ÿLf…x†cosnx LdxˆLan;ZL ÿLf…x†sinnx LdxˆLbn;ZL ÿLf…x†dxˆLa0: The required result follows on dividing both sides of Eq. (4.11) by L. 153PARSEVAL’S IDENTITY Parseval’s identity shows a relation between the average of the square of f…x† and the coecients in the Fourier series for f…x†: the average of ff…x†g2isRL ÿL‰f…x†Š2dx=2L; the average of ( a0=2†is (a0=2†2; the average of ( ancosnx†isa2 n=2; the average of ( bnsinnx†isb2 n=2. Example 4.4 Expand f…x†ˆx;0<x<2, in a half-range cosine series, then write Parseval’s identity corresponding to this Fourier cosine series. Solution: We first extend the definition of f…x†to that of the even function of period 4 shown in Fig. 4.7. Then 2 Lˆ4;Lˆ2. Thus bnˆ0 and anˆ2 LZL 0f…x†cosnx Ldxˆ2 2Z2 0f…x†cosnx 2dx ˆx2 nsinnx 2 ÿ1ÿ4 n22cosnx 2 2 0 ˆÿ4 n22cosnÿ1 …† ifn6ˆ0: Ifnˆ0, a0ˆZL 0xdxˆ2: Then f…x†ˆ1‡X1 nˆ14 n22cosnÿ1 …† cosnx 2: We now write Parseval’s identity. We first compute the average of ‰f…x†Š2: the average of ‰f…x†Š2ˆ1 2Z2 ÿ2f…x†fg2dxˆ12Z2 ÿ2x2dxˆ8 3; 154FOURIER SERIES AND INTEGRALS Figure 4.7. then the average a2 0 2‡X1 nˆ1a2n‡b2nÿ ˆ…2†2 2‡X1 nˆ116 n44cosnÿ1 …†2: Parseval’s identity now becomes 8 3ˆ2‡64 41 14‡1 34‡1 54‡ ; or 1 14‡1 34‡1 54‡ˆ4 96 which shows that we can use Parseval’s identity to find the sum of an infinite series. With the help of the above result, we can find the sum Sof the following series: 1 14‡1 24‡1 34‡1 44‡‡1 n4‡ : Sˆ1 14‡1 24‡1 34‡1 44‡ˆ1 14‡1 34‡1 54‡ ‡1 24‡1 44‡1 64‡ ˆ1 14‡1 34‡1 54‡ ‡1 241 14‡1 24‡1 34‡1 44‡ ˆ4 96‡S 16 from which we find Sˆ4=90. /C65lternati/C118e forms of Fourier series Up to this point the Fourier series of a function has been written as an infinite series of sines and cosines, Eq. (4.2): f…x†ˆa0 2‡X1 nˆ1ancosnx L‡bnsinnx L : This can be converted into other forms. In this section, we just discuss two alter- native forms. Let us first write, with =Lˆ ancosn x‡bnsinn xˆ a2n‡b2nqan a2n‡b2n/C112 cosn x‡bn a2n‡b2n/C112 sinn x/C32! : 155ALTERNATIVE FORMS OF FOURIER SERIES Now let (see Fig. 4.8) cosnˆan a2n‡b2n/C112 ;sinnˆbn a2n‡b2n/C112 ;sonˆtanÿ1bn an ; Cnˆ a2n‡b2nq ;C0ˆ1 2a0; then we have the trigonometric identity ancosn x‡bnsinn xˆCncosn xÿn …† ; and accordingly the Fourier series becomes f…x†ˆC0‡X1 nˆ1Cncosn xÿn …† : …4:12† In this new form, the Fourier series represents a periodic function as a sum of sinusoidal components having di/C128erent frequencies. The sinusoidal component of frequency n is called the nth harmonic of the periodic function. The first har- monic is commonly called the fundamental component. The angles nand the coecients Cnare known as the phase angle and amplitude. Using Euler’s identities eiˆcosisinwhere i2ˆÿ1, the Fourier series forf…x†can be converted into complex form f…x†ˆX1 nˆÿ1cneinx=L; …4:13a† where cnˆan/C7ibn ˆ1 2LZL ÿLf…x†eÿinx=Ldx;forn/C620: …4:13b† Eq. (4.13a) is obtained on the understanding that the Dirichlet conditions are satisfied and that f…x†is continuous at x.I ff…x†is discontinuous at x, the left hand side of Eq. (4.13a) should be replaced by ‰f…x‡0†‡f…xÿ0†Š=2. The exponential form (4.13a) can be considered as a basic form in its own right: it is not obtained by transformation from the trigonometric form, rather it is 156FOURIER SERIES AND INTEGRALS Figure 4.8. constructed directly from the given function. Furthermore, in the complex repre- sentation defined by Eqs. (4.13a) and (4.13b), a certain symmetry between theexpressions for a function and for its Fourier coecients is evident. In fact the expressions (4.13a) and (4.13b) are of essentially the same structure, as the follow- ing correlation reveals: x/C24L;f…x†/C24c nc…n†;einx=L/C24eÿinx=L;X1 nˆÿ1…†/C241 2LZL ÿL…†dx: This duality is worthy of note, and as our development proceeds to the Fourierintegral, it will become more striking and fundamental. Integration and di/C128erentiation of a Fourier series The Fourier series of a function f…x†may always be integrated term-by-term to give a new series which converges to the integral of f…x†.I ff…x†is a continuous function of xfor all x, and is periodic (of period 2 ) outside the interval ÿ<x<, then term-by-term di/C128erentiation of the Fourier series of f…x† leads to the Fourier series of f 0…x†, provided f0…x†satisfies Dirichlet’s conditions. /C86ibrating strings /C84he equation of motion of transverse vibration There are numerous applications of Fourier series to solutions of boundary value problems. Here we consider one of them, namely vibrating strings. Let a string of length Lbe held fixed between two points (0, 0) and ( L, 0) on the x-axis, and then given a transverse displacement parallel to the y-axis. Its subsequent motion, with no external forces acting on it, is to be considered; this is described by finding thedisplacement yas a function of xandt(if we consider only vibration in one plane, and take the xyplane as the plane of vibration). We will assume that /C26, the mass per unit length is uniform over the entire length of the string, and that the string is perfectly flexible, so that it can transmit tension but not bending or shearing forces. As the string is drawn aside from its position of rest along the x-axis, the resulting increase in length causes an increase in tension, denoted by P. This tension at any point along the string is always in the direction of the tangent tothe string at that point. As shown in Fig. 4.9, a force P…x†Aacts at the left hand side of an element ds, and a force P…x‡dx†Aacts at the right hand side, where A is the cross-sectional area of the string. If is the inclination to the horizontal, then /C70 x/C129APcos… ‡d †ÿAPcos ; /C70y/C129APsin… ‡d †ÿAPsin : 157INTEGRATION AND DIFFERENTIATION OF A FOURIER SERIES We limit the displacement to small values, so that we may set cos ˆ1ÿ 2=2;sin /C129 /C129tan ˆdy=dx; then /C70yˆAPdy dx x‡dxÿdy dx x ˆAPd2y dx2dx: Using Newton’s second law, the equation of motion of transverse vibration of the element becomes /C26Adx/C642y /C64t2ˆAP/C642y /C64x2dx; or/C642y /C64x2ˆ1 /C1182/C642y /C64t2;/C118ˆ  P=/C26/C112 : Thus the transverse displacement of the string satisfies the partial di/C128erential wave equation /C642y /C64x2ˆ1 /C1182/C642y /C64t2;0<x<L;t/C620 …4:14† with the following boundary conditions: y…0;t†ˆy…L;t†ˆ0;/C64y=/C64tˆ0; y…x;0†ˆf…x†; where f…x†describes the initial shape (position) of the string, and/C118is the velocity of propagation of the wave along the string. Solution of the /C119ave equation To solve this boundary value problem, let us try the method of separation vari- ables: y…x;t†ˆX…x†T…t†: …4:15† Substituting this into Eq. (4.14) yields …1=X†…d2X=dx2†ˆ… 1=/C1182T†…d2T=dt2†. 158FOURIER SERIES AND INTEGRALS Figure 4.9. A vibrating string. Since the left hand side is a function of xonly and the right hand side is a function of time only, they must be equal to a common separation constant, which we will callÿ2. Then we have d2X=dx2ˆÿ2X;X…0†ˆX…L†ˆ0 …4:16a† and d2T=dt2ˆÿ2/C1182Td T =dtˆ0a t tˆ0: …4:16b† Both of these equations are typical eigenvalue problems: we have a di/C128erentialequation containing a parameter , and we seek solutions satisfying certain boundary conditions. If there are special values of for which non-trivial solu- tions exist, we call these eigenvalues, and the corresponding solutions eigensolu-tions or eigenfunctions. The general solution of Eq. (4.16a) can be written as X…x†ˆA 1sin…x†‡B1cos…x†: Applying the boundary conditions X…0†ˆ0/C41B1ˆ0; and X…L†ˆ0/C41A1sin…L†ˆ0 A1ˆ0 is the trivial solution Xˆ0 (so yˆ0); hence we must have sin …L†ˆ0, that is, Lˆn;nˆ1;2;...; and we obtain a series of eigenvalues nˆn=L;nˆ1;2;... and the corresponding eigenfunctions Xn…x†ˆsin…n=L†x;nˆ1;2;...: To solve Eq. (4.16b) for T…t†we must use one of the values nfound above. The general solution is of the form T…t†ˆA2cos…n/C118t†‡B2sin…n/C118t†: The boundary condition leads to B2ˆ0. The general solution of Eq. (4.14) is hence a linear superposition of the solu- tions of the form y…x;t†ˆX1 nˆ1Ansin…nx=L†cos…n/C118t=L†; …4:17† 159VIBRATING STRINGS theAnare as yet undetermined constants. To find An, we use the boundary condition y…x;t†ˆf…x†attˆ0, so that Eq. (4.17) reduces to f…x†ˆX1 nˆ1Ansin…nx=L†: Do you recognize the infinite series on the right hand side/C63 It is a Fourier sine series. To find An, multiply both sides by sin( mx=L) and then integrate with respect to xfrom 0 to Land we obtain Amˆ2 LZL 0f…x†sin…mx=L†dx; mˆ1;2;... where we have used the relation ZL 0sin…mx=L†sin…nx=L†dxˆL 2mn: Eq. (4.17) now gives y…x;t†ˆX1 nˆ12 LZL 0f…x†sinnx Ldx sinnx Lcosn/C118t L: …4:18† The terms in this series represent the natural modes of vibration. The frequency of the nth normal mode fnis obtained from the term involving cos …n/C118t=L†and is given by 2fnˆn/C118=Lorfnˆn/C118=2L: All frequencies are integer multiples of the lowest frequency f1.W ec a l l f1the fundamental frequency or first harmonic, and f2andf3the second and third harmonics (or first and second overtones) and so on. /C82L/C67 circuit Another good example of application of Fourier series is an /C82L/C67 circuit driven by a variable voltage /C69…t†which is periodic but not necessarily sinusoidal (see Fig. 4.10). We want to find the current I…t†flowing in the circuit at time t. According to Kirchho/C128 ’s second law for circuits, the impressed voltage /C69…t† equals the sum of the voltage drops across the circuit components. That is, LdI dt‡RI‡/C81 Cˆ/C69…t†; where /C81is the total charge in the capacitor /C67. But Iˆd/C81=dt, thus di/C128erentiating the above di/C128erential equation once we obtain Ld2I dt2‡RdI dt‡1 CIˆd/C69 dt: 160FOURIER SERIES AND INTEGRALS Under steady-state conditions the current I…t†is also periodic, with the same period Pas for /C69…t†. Let us assume that both /C69…t†andI…t†possess Fourier expansions and let us write them in their complex forms: /C69…t†ˆX1 nˆÿ1/C69nein/C33t; I…t†ˆX1 nˆÿ1cnein/C33t…/C33ˆ2=P†: Furthermore, we assume that the series can be di/C128erentiated term by term. Thus d/C69 dtˆX1 nˆÿ1in/C33/C69nein/C33t;dI dtˆX1 nˆÿ1in/C33cnein/C33t;d2I dt2ˆX1 nˆÿ1…ÿn2/C332†cnein/C33t: Substituting these into the last (second-order) di/C128erential equation and equating the coecients with the same exponential ein t, we obtain ÿn2/C332L‡in/C33R‡1=Cÿ cnˆin/C33/C69n: Solving for cn cnˆin/C33=L ‰1=CL…†2ÿn2/C332ЇiR=L…† n/C33/C69n: Note that 1/ L/C67is the natural frequency of the circuit and /C82/C47L is the attenuation factor of the circuit. The Fourier coecients for /C69…t†are given by /C69nˆ1 PZP=2 ÿP=2/C69…t†eÿin/C33tdt: The current I…t†in the circuit is given by I…t†ˆX1 nˆÿ1cnein/C33t: 161/C82L/C67 CIRCUIT Figure 4.10. The /C82L/C67 circuit. /C79rthogonal functions Many of the properties of Fourier series considered above depend on orthogonal properties of sine and cosine functions ZL 0sinmx Lsinnx Ldxˆ0;ZL 0cosmx Lcosnx Ldxˆ0…m6ˆn†: In this section we seek to generalize this orthogonal property. To do so we firstrecall some elementary properties of real vectors in three-dimensional space. Two vectors /C65and/C66are called orthogonal if /C65/C66ˆ0. Although not geome- trically or physically obvious, we generalize these ideas to think of a function, say A…x†, as being an infinite-dimensional vector (a vector with an infinity of compo- nents), the value of each component being specified by substituting a particularvalue of xtaken from some interval ( a,b), and two functions, A…x†andB…x†are orthogonal in ( a,b)i f Z b aA…x†B…x†dxˆ0: …4:19† The left-side of Eq. (4.19) is called the scalar product of A…x†andB…x†and denoted by, in the Dirac bracket notation, hA…x†jB…x†i. The first factor in the bracket notation is referred to as the bra and the second factor as the ket, sotogether they comprise the bracket. A vector /C65is called a unit vector or normalized vector if its magnitude is unity: /C65/C65ˆA 2ˆ1. Extending this concept, we say that the function A…x†is normal or normalized in ( a,b)i f hA…x†jA…x†i ˆZb aA…x†A…x†dxˆ1: …4:20† If we have a set of functions ’i…x†;iˆ1;2;3;...;having the properties ’m…x† hj ’n…x†i ˆZb a’m…x†’n…x†dxˆmn; …4:20a† where nmis the Kronecker delta symbol, we then call such a set of functions an orthonormal set in ( a,b). For example, the set of functions ’m…x†ˆ …2=†1=2sin…mx†;mˆ1;2;3;...is an orthonormal set in the interval 0 x. Just as in three-dimensional vector space, any vector /C65can be expanded in the form /C65ˆA1^e1‡A2^e2‡A3^e3, we can consider a set of orthonormal functions ’i as base vectors and expand a function f…x†in terms of them, that is, f…x†ˆX1 nˆ1cn’n…x†axb; …4:21† 162FOURIER SERIES AND INTEGRALS the series on the right hand side is called an orthonormal series; such series are generalizations of Fourier series. Assuming that the series on the right convergestof…x†, we can then multiply both sides by ’ m…x†and integrate both sides from a tobto obtain cmˆhf…x†j’m…x†i ˆZb af…x†’m…x†dx; …4:21a† cmcan be called the generalized Fourier coecients. Multiple Fourier series A Fourier expansion of a function of two or three variables is often very useful inmany applications. Let us consider the case of a function of two variables, say f…x;y†. For example, we can expand f…x;y†into a double Fourier sine series f…x;y†ˆX 1 mˆ1X1 nˆ1Bmnsinmx L1sinny L2; …4:22† where Bmnˆ4 L1L2ZL1 0ZL2 0f…x;y†sinmx L1sinny L2dxdy : …4:22a† Similar expansions can be made for cosine series and for series having both sines and cosines. To obtain the coecients Bmn, let us rewrite f…x;y†as f…x;y†ˆX1 mˆ1Cmsinmx L1; …4:23† where CmˆX1 nˆ1Bmnsinny L2: …4:23a† Now we can consider Eq. (4.23) as a Fourier series in which yis kept constant so that the Fourier coecients Cmare given by Cmˆ2 L1ZL1 0f…x;y†sinmx L1dx: …4:24† On noting that Cmis a function of y, we see that Eq. (4.23a) can be considered as a Fourier series for which the coecients Bmnare given by Bmnˆ2 L2ZL2 0Cmsinny L2dy: 163MULTIPLE FOURIER SERIES Substituting Eq. (4.24) for Cminto the above equation, we see that Bmnis given by Eq. (4.22a). Similar results can be obtained for cosine series or for series containing both sines and cosines. Furthermore, these ideas can be generalized to triple Fourier series, etc. They are very useful in solving, for example, wave propagation and heat conduction problems in two or three dimensions. Because they lie outside of the scope of this book, we have to omit these interesting applications. Fourier integrals and Fourier transforms The properties of Fourier series that we have thus far developed are adequate forhandling the expansion of any periodic function that satisfies the Dirichlet con- ditions. But many problems in physics and engineering do not involve periodic functions, and it is therefore desirable to generalize the Fourier series method toinclude non-periodic functions. A non-periodic function can be considered as a limit of a given periodic function whose period becomes infinite, as shown in Examples 4.5 and 4.6. Example 4.5 Consider the periodic functions f L…x† fL…x†ˆ0 when ÿL=2<x<ÿ1 1 when ÿ1<x<1 0 when 1 <x<L=28 >< >:; 164FOURIER SERIES AND INTEGRALS Figure 4.11. Square wave function: …a†Lˆ4;…b†Lˆ8;…c†L!1 . which has period L/C622. Fig. 4.11( a) shows the function when Lˆ4. If Lis increased to 8, the function looks like the one shown in Fig. 4.11( b). As L!1 we obtain a non-periodic function f…x†, as shown in Fig. 4.11( c): f…x†ˆ1ÿ1<x<1 0 otherwise/C26 : Example 4.6 Consider the periodic function /C103L…x†(Fig. 4.12( a)): /C103L…x†ˆeÿjxjwhen ÿL=2<x<L=2: AsL!1 we obtain a non-periodic function /C103…x†/C58/C103…x†ˆlimL!1/C103L…x†(Fig. 4.12( b)). By investigating the limit that is approached by a Fourier series as the period of the given function becomes infinite, a suitable representation for non-periodic functions can perhaps be obtained. To this end, let us write the Fourier series representing a periodic function f…x†in complex form: f…x†ˆX1 nˆÿ1cnei/C33x; …4:25† cnˆ1 2LZL ÿLf…x†eÿi/C33xdx …4:26† where /C33denotes n=L /C33ˆn L;npositive or negative : …4:27† The transition L!1 is a little tricky since cnapparently approaches zero, but these coecients should not approach zero. We can ask for help from Eq. (4.27), from which we have /C33ˆ…=L†n; 165FOURIER INTEGRALS AND FOURIER TRANSFORMS Figure 4.12. Sawtooth wave functions: …a†ÿL=2<x<L=2;…b†L!1 . and the ‘adjacent’ values of /C33are obtained by setting nˆ1, which corresponds to …L=†/C33ˆ1: Then we can multiply each term of the Fourier series by …L=†/C33and obtain f…x†ˆX1 nˆÿ1L cn ei/C33x/C33; where L cnˆ1 2ZL ÿLf…x†eÿi/C33xdx: The troublesome factor 1 =Lhas disappeared. Switching completely to the /C33 notation and writing …L=†cnˆcL…/C33†, we obtain cL…/C33†ˆ1 2ZL ÿLf…x†eÿi/C33xdx and f…x†ˆX1 L/C33=ˆÿ1cL…/C33†ei/C33x/C33: In the limit as L!1 , the /C33s are distributed continuously instead of discretely, /C33!d/C33and this sum is exactly the definition of an integral. Thus the last equations become c…/C33†ˆlim L!1cL…/C33†ˆ1 2Z1 ÿ1f…x†eÿi/C33xdx …4:28† and f…x†ˆZ1 ÿ1c…/C33†ei/C33xd/C33: …4:29† This set of formulas is known as the Fourier transformation, in somewhat di/C128er- ent form. It is easy to put them in a symmetrical form by defining /C103…/C33†ˆ  2p c…ÿ/C33†; then Eqs. (4.28) and (4.29) take the symmetrical form /C103…/C33†ˆ1  2pZ1 ÿ1f…x0†eÿi/C33x0dx0; …4:30† f…x†ˆ1 2pZ 1 ÿ1/C103…/C33†ei/C33xd/C33: …4:31† 166FOURIER SERIES AND INTEGRALS The function /C103…/C33†is called the Fourier transform of f…x†and is written /C103…/C33†ˆ/C70ff…x†g. Eq. (4.31) is the inverse Fourier transform of /C103…/C33†and is written f…x†ˆ/C70ÿ1f/C103…/C33†g; sometimes it is also called the Fourier integral representation off…x†. The exponential function eÿi/C33xis sometimes called the kernel of trans- formation. It is clear that /C103…/C33†is defined only if f…x†satisfies certain restrictions. For instance, f…x†should be integrable in some finite region. In practice, this means thatf…x†has, at worst, jump discontinuities or mild infinite discontinuities. Also, the integral should converge at infinity. This would require that f…x†!0a s x! 1 . A very common sucient condition is the requirement that f…x†is absolutely integrable. That is, the integral Z1 ÿ1f…x†jj dx exists. Since jf…x†eÿi/C33xjˆjf…x†j, it follows that the integral for /C103…/C33†is absolutely convergent; therefore it is convergent. It is obvious that /C103…/C33†is, in general, a complex function of the real variable /C33. So if f…x†is real, then /C103…ÿ/C33†ˆ/C103/C42…/C33†: There are two immediate corollaries to this property: (1)f…x†is even, /C103…/C33†is real; (2) if f…x†is odd, /C103…/C33†is purely imaginary. Other, less symmetrical forms of the Fourier integral can be obtained by working directly with the sine and cosine series, instead of with the exponential functions. Example 4.7 Consider the Gaussian probability function f…x†ˆ/C78eÿ x2, where /C78and are constant. Find its Fourier transform /C103…/C33†, then graph f…x†and/C103…/C33†. Solution: Its Fourier transform is given by /C103…/C33†ˆ1  2pZ1 ÿ1f…x†eÿi/C33xdxˆ/C78 2pZ 1 ÿ1eÿ x2eÿi/C33xdx: This integral can be simplified by a change of variable. First, we note that ÿ x2ÿi/C33xˆÿ … x  p‡i/C33=2  p†2ÿ/C332=4 ; and then make the change of variable x  p‡i/C33=2  pˆuto obtain /C103…/C33†ˆ /C78 2 p eÿ/C332=4 Z1 ÿ1eÿu2duˆ/C782 peÿ/C332=4 : 167FOURIER INTEGRALS AND FOURIER TRANSFORMS It is easy to see that /C103…/C33†is also a Gaussian probability function with a peak at the origin, monotonically decreasing as /C33! 1 . Furthermore, for large ,f…x†is sharply peaked but /C103…/C33†is flattened, and vice versa as shown in Fig. 4.13. It is interesting to note that this is a general feature of Fourier transforms. We shall see later that in quantum mechanical applications it is related to the Heisenberg uncertainty principle. The original function f…x†can be retrieved from Eq. (4.31) which takes the form 1  2pZ1 ÿ1/C103…/C33†ei/C33xd/C33ˆ1 2p /C78 2 pZ1 ÿ1eÿ/C332=4 ei/C33xd/C33 ˆ1  2p/C78 2 pZ1 ÿ1eÿ 0/C332eÿi/C33x0d/C33 in which we have set 0ˆ1=4 , and x0ˆÿx. The last integral can be evaluated by the same technique, and we finally find 1  2pZ1 ÿ1/C103…/C33†ei/C33xd/C33ˆ1 2p /C78 2 pZ1 ÿ1eÿ 0/C332eÿi/C33x0d/C33 ˆ/C782 p 2 p eÿ x2 ˆ/C78eÿ x2ˆf…x†: Example 4.8 Given the box function which can represent a single pulse f…x†ˆ1jxja 0xjj/C62a/C26 find the Fourier transform of f…x†,/C103…/C33†; then graph f…x†and/C103…/C33†foraˆ3. 168FOURIER SERIES AND INTEGRALS Figure 4.13. Gaussian probability function: …a†large ;…b†small . Solution: The Fourier transform of f…x†is, as shown in Fig. 4.14, /C103…/C33†ˆ1  2pZ1 ÿ1f…x0†eÿi/C33x0dx0ˆ1 2pZ a ÿa…1†eÿi/C33x0dx0ˆ1 2p eÿi/C33x0 ÿi/C33a ÿa/C12/C12/C12/C12/C12 ˆ 2 /C114 sin/C33a /C33;/C336ˆ0: For/C33ˆ0, we obtain /C103…/C33†ˆ 2=/C112 a. The Fourier integral representation of f…x†is f…x†ˆ1  2pZ1 ÿ1/C103…/C33†ei/C33xd/C33ˆ1 2Z1 ÿ12 sin /C33a /C33ei/C33xd/C33: Now Z1 ÿ1sin/C33a /C33ei/C33xd/C33ˆZ1 ÿ1sin/C33acos/C33x /C33d/C33‡iZ1 ÿ1sin/C33asin/C33x /C33d/C33: The integrand in the second integral is odd and so the integral is zero. Thus we have f…x†ˆ1  2pZ1 ÿ1/C103…/C33†ei/C33xd/C33ˆ1 Z1 ÿ1sin/C33acos/C33x /C33d/C33ˆ2 Z1 0sin/C33acos/C33x /C33d/C33; the last step follows since the integrand is an even function of /C33. It is very dicult to evaluate the last integral. But a known property of f…x†will help us. We know that f…x†is equal to 1 for jxja, and equal to 0 for jxj/C62a. Thus we can write 2 Z1 0sin/C33acos/C33x /C33d/C33ˆ1jxja 0jxj/C62a/C26 169FOURIER INTEGRALS AND FOURIER TRANSFORMS Figure 4.14. The box function. Just as in Fourier series expansion, we also expect to observe Gibb’s phenomenon in the case of Fourier integrals. Approximations to the Fourier integral are obtained by replacing 1by : Z 0sin/C33cos/C33x /C33d/C33; where we have set aˆ1. Fig. 4.15 shows oscillations near the points of disconti- nuity of f…x†. We might expect these oscillations to disappear as !1 , but they are just shifted closer to the points xˆ1. Example 4.9 Consider now a harmonic wave of frequency /C330,ei/C330t, which is chopped to a life- time of 2 Tseconds (Fig. 4.16( a)): f…t†ˆei/C330tÿTtT 0 jtj/C620:( The chopping process will introduce many new frequencies in varying amounts,given by the Fourier transform. Then we have, according to Eq. (4.30), /C103…/C33†ˆ… 2† ÿ1=2ZT ÿTei/C330teÿi/C33tdtˆ…2†ÿ1=2ZT ÿTei…/C330ÿ/C33†tdt ˆ…2†ÿ1=2ei…/C330ÿ/C33†t i…/C330ÿ/C33†/C12/C12/C12/C12T ÿTˆ…2=†1=2Tsin…/C330ÿ/C33†T …/C330ÿ/C33†T: This function is plotted schematically in Fig. 4.16( b). (Note that limx!0…sinx=x†ˆ1.) The most striking aspect of this graph is that, although the principal contribution comes from the frequencies in the neighborhood of /C330, an infinite number of frequencies are presented. Nature provides an example of this kind of chopping in the emission of photons during electronic and nucleartransitions in atoms. The light emitted from an atom consists of regular vibrationsthat last for a finite time of the order of 10 ÿ9s or longer. When light is examined by a spectroscope (which measures the wavelengths and, hence, the frequencies) we find that there is an irreducible minimum frequency spread for each spectrum line. This is known as the natural line width of the radiation. The relative percentage of frequencies, other than the basic one, present depends on the shape of the pulse, and the spread of frequencies depends on 170FOURIER SERIES AND INTEGRALS Figure 4.15. The Gibb’s phenomenon. the time Tof the duration of the pulse. As Tbecomes larger the central peak becomes higher and the width /C33…ˆ2=T†becomes smaller. Considering only the spread of frequencies in the central peak we have /C33ˆ2=T;orTˆ1: Multiplying by the Planck constant hand replacing Tbyt, we have the relation t/C69ˆ/C104: …4:32† A wave train that lasts a finite time also has a finite extension in space. Thus the radiation emitted by an atom in 10ÿ9s has an extension equal to 3 10810ÿ9ˆ 310ÿ1m. A Fourier analysis of this pulse in the space domain will yield a graph identical to Fig. 4.11( b), with the wave numbers clustered around k0…ˆ2= 0ˆ/C330=/C118†. If the wave train is of length 2 a, the spread in wave number will be given by akˆ2, as shown below. This time we are chopping an infinite plane wave front with a shutter such that the length of the packet is 2 a, where 2aˆ2/C118T, and 2 Tis the time interval that the shutter is open. Thus /C32…x†ˆeik0x;ÿaxa 0; jxj/C62a:( Then /C30…k†ˆ… 2†ÿ1=2Z1 ÿ1/C32…x†eÿikxdxˆ…2†ÿ1=2Za ÿa/C32…x†eÿikxdx ˆ…2=†1=2asin…k0ÿk†a …k0ÿk†a: This function is plotted in Fig. 4.17: it is identical to Fig. 4.16( b), but here it is the wave vector (or the momentum) that takes on a spread of values around k0. The breadth of the central peak is kˆ2=a,o rakˆ2. 171FOURIER INTEGRALS AND FOURIER TRANSFORMS Figure 4.16. ( a) A chopped harmonic wave ei/C330tthat lasts a finite time 2 T.…b† Fourier transform of e…i/C330t†;jtj<T, and 0 otherwise. Fourier sine and cosine transforms Iff…x†is an odd function, the Fourier transforms reduce to /C103…/C33†ˆ 2 /C114Z1 0f…x0†sin/C33x0dx0; f…x†ˆ 2 /C114Z1 0/C103…/C33†sin/C33xd/C33: …4:33a† Similarly, if f…x†is an even function, then we have Fourier cosine transforma- tions: /C103…/C33†ˆ 2 /C114Z1 0f…x0†cos/C33x0dx0; f…x†ˆ 2 /C114Z1 0/C103…/C33†cos/C33xd/C33:…4:33b† To demonstrate these results, we first expand the exponential function on the right hand side of Eq. (4.30) /C103…/C33†ˆ1  2pZ1 ÿ1f…x0†eÿi/C33x0dx0 ˆ1 2pZ 1 ÿ1f…x0†cos/C33x0dx0ÿi 2pZ 1 ÿ1f…x0†sin/C33x0dx0: Iff…x†is even, then f…x†cos/C33xis even and f…x†sin/C33xis odd. Thus the second integral on the right hand side of the last equation is zero and we have /C103…/C33†ˆ1 2pZ 1 ÿ1f…x0†cos/C33x0dx0ˆ 2 /C114Z1 0f…x0†cos/C33x0dx0; /C103…/C33†is an even function, since /C103…ÿ/C33†ˆ/C103…/C33†. Next from Eq. (4.31) we have f…x†ˆ1  2pZ1 ÿ1/C103…/C33†ei/C33xd/C33 ˆ1  2pZ1 ÿ1/C103…/C33†cos/C33xd/C33‡i 2pZ 1 ÿ1/C103…/C33†sin/C33xd/C33: 172FOURIER SERIES AND INTEGRALS Figure 4.17. Fourier transform of eikx;jxja: Since /C103…/C33†is even, so /C103…/C33†sin/C33xis odd and the second integral on the right hand side of the last equation is zero, and we have f…x†ˆ1  2pZ1 ÿ1/C103…/C33†cos/C33xd/C33ˆ 2 /C114Z1 0/C103…/C33†cos/C33xd/C33: Similarly, we can prove Fourier sine transforms by replacing the cosine by the sine. /C72eisenberg/C39s uncertaint/C121 principle We have demonstrated in above examples that if f…x†is sharply peaked, then /C103…/C33† is flattened, and vice versa. This is a general feature in the theory of Fourier transforms and has important consequences for all instances of wave propaga- tion. In electronics we understand now why we use a wide-band amplification in order to reproduce a sharp pulse without distortion. In quantum mechanical applications this general feature of the theory of Fourier transforms is related to the Heisenberg uncertainty principle. We sawin Example 4.9 that the spread of the Fourier transform in kspace ( k) times its spread in coordinate space ( a) is equal to 2 …ak/C1292†. This result is of special importance because of the connection between values of kand momentum /C112/C58/C112ˆpk(where pis the Planck constant hdivided by 2 ). A particle localized in space must be represented by a superposition of waves with di/C128erent momenta. As a result, the position and momentum of a particle cannot be measured simul/C45 taneously with infinite precision; the product of ‘uncertainty in the position deter- mination’ and ‘uncertainty in the momentum determination’ is governed by the relation x/C112/C129/C104…apk/C1292pˆ/C104,o r x/C112/C129/C104;xˆa†. This statement is called Heisenberg’s uncertainty principle. If position is known better, knowledge of the momentum must be unavoidably reduced proportionally, and vice versa. A complete knowledge of one, say k(and so p), is possible only when there is complete ignorance of the other. We can see this in physical terms. A wavewith a unique value of kis infinitely long. A particle represented by an infinitely long wave (a free particle) cannot have a definite position, since the particle can beanywhere along its length. Hence the position uncertainty is infinite in order that the uncertainty in kis zero. Equation (4.32) represents Heisenberg’s uncertainty principle in a di/C128erent form. It states that we cannot know with infinite precision the exact energy of aquantum system at every moment in time. In order to measure the energy of a quantum system with good accuracy, one must carry out such a measurement for a suciently long time. In other words, if the dynamical state exists only for a time of order t, then the energy of the state cannot be defined to a precision better than /C104=t. 173HEISENBERG’S UNCERTAINTY PRINCIPLE We should not look upon the uncertainty principle as being merely an unfor- tunate limitation on our ability to know nature with infinite precision. We can use it to our advantage. For example, when combining the time–energy uncertainty relation with Einstein’s mass–energy relation ( /C69ˆmc2) we obtain the relation mt/C129/C104=c2. This result is very useful in our quest to understand the universe, in particular, the origin of matter. /C87a/C118e pac/C107ets and group /C118elocit/C121 Energy (that is, a signal or information) is transmitted by groups of waves, not asingle wave. Phase velocity may be greater than the speed of light c, ‘group velocity’ is always less than c. The wave groups with which energy is transmitted from place to place are called wave packets. Let us first consider a simple casewhere we have two waves ’ 1and’2: each has the same amplitude but di/C128ers slightly in frequency and wavelength, ’1…x;t†ˆAcos…/C33tÿkx†; ’2…x;t†ˆAcos‰…/C33‡/C33†tÿ…k‡k†xŠ; where /C33/C28/C33and k/C28k. Each represents a pure sinusoidal wave extending to infinite along the x-axis. Together they give a resultant wave ’ˆ’1‡’2 ˆAcos…/C33tÿkx†‡cos‰…/C33‡/C33†tÿ…k‡k†xŠ fg : Using the trigonometrical identity cosA‡cosBˆ2 cosA‡B 2cosAÿB 2; we can rewrite ’as ’ˆ2 cos2/C33tÿ2kx‡/C33tÿkx 2cosÿ/C33t‡kx 2 ˆ2 cos1 2…/C33tÿkx†cos…/C33tÿkx†: This represents an oscillation of the original frequency /C33, but with a modulated amplitude as shown in Fig. 4.18. A given segment of the wave system, such as AB, can be regarded as a ‘wave packet’ and moves with a velocity /C118/C103(not yet deter- mined). This segment contains a large number of oscillations of the primary wave that moves with the velocity /C118. And the velocity /C118/C103with which the modulated amplitude propagates is called the group velocity and can be determined by therequirement that the phase of the modulated amplitude be constant. Thus /C118 /C103ˆdx=dtˆ/C33=k!d/C33=dk: 174FOURIER SERIES AND INTEGRALS The modulation of the wave is repeated indefinitely in the case of superposition of two almost equal waves. We now use the Fourier technique to demonstrate thatany isolated packet of oscillatory disturbance of frequency /C33can be described in terms of a combination of infinite trains of frequencies distributed around /C33. Let us first superpose a system of nwaves /C32…x;t†ˆX n jˆ1Ajei…kjxÿ/C33jt†; where Ajdenotes the amplitudes of the individual waves. As napproaches infinity, the frequencies become continuously distributed. Thus we can replace the sum- mation with an integration, and obtain /C32…x;t†ˆZ1 ÿ1A…k†ei…kxÿ/C33t†dk; …4:34† the amplitude A…k†is often called the distribution function of the wave. For /C32…x;t†to represent a wave packet traveling with a characteristic group velocity, it is necessary that the range of propagation vectors included in the superposition be fairly small. Thus, we assume that the amplitude A…k†6 ˆ0 only for a small range of values about a particular k0ofk: A…k†6 ˆ0;k0ÿ/C34<k<k0‡/C34; /C34 /C28k0: The behavior in time of the wave packet is determined by the way in which the angular frequency /C33depends upon the wave number k/C58/C33ˆ/C33…k†, known as the law of dispersion. If /C33varies slowly with k, then /C33…k†can be expanded in a power series about k0: /C33…k†ˆ/C33…k0†‡d/C33 dk/C12/C12/C12/C12 0…kÿk0†‡ˆ /C330‡/C330…kÿk0†‡O…kÿk0†2/C104/C105 ; where /C330ˆ/C33…k0†;and /C330ˆd/C33 dk/C12/C12/C12/C12 0 175WAVE PACKETS AND GROUP VELOCITY Figure 4.18. Superposition of two waves. and the subscript zero means ‘evaluated’ at kˆk0. Now the argument of the exponential in Eq. (4.34) can be rewritten as /C33tÿkxˆ…/C330tÿk0x†‡/C330…kÿk0†tÿ…kÿk0†x ˆ…/C330tÿk0x†‡…kÿk0†…/C330tÿx† and Eq. (4.34) becomes /C32…x;t†ˆexp‰i…k0xÿ/C330t†ŠZk0‡/C34 k0ÿ/C34A…k†exp‰i…kÿk0†…xÿ/C330t†Šdk: …4:35† If we take kÿk0as the new integration variable yand assume A…k†to be a slowly varying function of kin the integration interval 2 /C34, then Eq. (4.35) becomes /C32…x;t†/C129exp‰i…k0xÿ/C330t†ŠZk0‡/C34 k0ÿ/C34A…k0‡y†exp‰i…xÿ/C330t†yŠdy: Integration, transformation, and the approximation A…k0‡y†/C129A…k0†lead to the result /C32…x;t†ˆB…x;t†exp‰i…k0xÿ/C330t†Š … 4:36† with B…x;t†ˆ2A…k0†sin‰k…xÿ/C330t†Š xÿ/C330t: …4:37† As the argument of the sine contains the small quantity k;B…x;t†varies slowly depending on time tand coordinate x. Therefore, we can regard B…x;t†as the small amplitude of an approximately monochromatic wave and k0xÿ/C330tas its phase. If we multiply the numerator and denominator on the right hand side of Eq. (4.37) by kand let zˆk…xÿ/C330t† then B…x;t†becomes B…x;t†ˆ2A…k0†ksinz z and we see that the variation in amplitude is determined by the factor sin ( z†=z. This has the properties lim z!0sinz zˆ1 for zˆ0 and sinz zˆ0 for zˆ;2;...: 176FOURIER SERIES AND INTEGRALS If we further increase the absolute value of z, the function sin ( z†=zruns alter- nately through maxima and minima, the function values of which are small com- pared with the principal maximum at zˆ0, and quickly converges to zero. Therefore, we can conclude that superposition generates a wave packet whoseamplitude is non-zero only in a finite region, and is described by sin ( z†=z(see Fig. 4.19). The modulating factor sin ( z†=zof the amplitude assumes the maximum value 1 asz!0. Recall that zˆk…xÿ/C33 0t), thus for zˆ0, we have xÿ/C330tˆ0; which means that the maximum of the amplitude is a plane propagating with velocity dx dtˆ/C330ˆd/C33 dk/C12/C12/C12/C12 0; that is, /C330is the group velocity, the velocity of the whole wave packet. The concept of a wave packet also plays an important role in quantum mechanics. The idea of associating a wave-like property with the electron and other material particles was first proposed by Louis Victor de Broglie (1892–1987) in 1925. His work was motivated by the mystery of the Bohr orbits. After Rutherford’s successful -particle scattering experiments, a planetary-type nuclear atom, with electrons orbiting around the nucleus, was in favor withmost physicists. But, according to classical electromagnetic theory, a charge undergoing continuous centripetal acceleration emits electromagnetic radiation 177WAVE PACKETS AND GROUP VELOCITY Figure 4.19. A wave packet. continuously and so the electron would lose energy continuously and it would spiral into the nucleus after just a fraction of a second. This does not occur.Furthermore, atoms do not radiate unless excited, and when radiation does occur its spectrum consists of discrete frequencies rather than the continuum of frequencies predicted by the classical electromagnetic theory. In 1913 Niels Bohr (1885–1962) proposed a theory which successfully explained the radiation spectra of the hydrogen atom. According to Bohr’s postulates, an atom can exist in certain allowed stationary states without radiation. Only when an electron makes a transition between two allowed stationary states, does it emit or absorb radiation. The possible stationary states are those in which the angular momen- tum of the electron about the nucleus is quantized, that is, m/C118rˆnp, where vis the speed of the electron in the nth orbit and r is its radius. Bohr didn’t clearly describe this quantum condition. De Broglie attempted to explain it by fitting astanding wave around the circumference of each orbit. Thus de Broglie proposed thatnˆ2r, where is the wavelength associated with the nth orbit. Combining this with Bohr’s quantum condition we immediately obtain ˆ /C104 m/C118ˆ/C104 /C112: De Broglie proposed that any material particle of total energy Eand momentum p is accompanied by a wave whose wavelength is given by ˆ/C104=/C112and whose frequency is given by the Planck formula /C23ˆ/C69=/C104. Today we call these waves de Broglie waves or matter waves. The physical nature of these matter waves was not clearly described by de Broglie, we shall not ask what these matter waves are – this is addressed in most textbooks on quantum mechanics. Let us ask just onequestion: what is the (phase) velocity of such a matter wave/C63 If we denote this velocity by u, then uˆ/C23ˆ/C69 /C112ˆ1 /C112 /C1122c2‡m2 0c4q ˆc 1‡…m0c=/C112†2q ˆc2 /C118/C112ˆm0/C118 1ÿ/C1182=c2/C112/C32! ; which shows that for a particle with m0/C620 the wave velocity uis always greater than c, the speed of light in a vacuum. Instead of individual waves, de Broglie suggested that we can think of particles inside a wave packet, synthesized from a number of individual waves of di/C128erent frequencies, with the entire packet travel- ing with the particle velocity /C118. De Broglie’s matter wave idea is one of the cornerstones of quantum mechanics. 178FOURIER SERIES AND INTEGRALS /C72eat conduction We now consider an application of Fourier integrals in classical physics. A semi- infinite thin bar ( x0), whose surface is insulated, has an initial temperature equal to f…x†. The temperature of the end xˆ0 is suddenly dropped to and maintained at zero. The problem is to find the temperature T…x;t†at any point xat time t. First we have to set up the boundary value problem for heat conduc- tion, and then seek the general solution that will give the temperature T…x;t†at any point xat time t. /C72ead conduction equation To establish the equation for heat conduction in a conducting medium we need first to find the heat flux (the amount of heat per unit area per unit time) across a surface. Suppose we have a flat sheet of thickness n, which has temperature Ton one side and T‡Ton the other side (Fig. 4.20). The heat flux which flows from the side of high temperature to the side of low temperature is directly proportionalto the di/C128erence in temperature Tand inversely proportional to the thickness n. That is, the heat flux from I to II is equal to ÿ/C75T n; where /C75, the constant of proportionality, is called the thermal conductivity of the conducting medium. The minus sign is due to the fact that if T/C620 the heat actually flows from II to I. In the limit of n!0, the heat flux across from II to I can be written ÿ/C75/C64T /C64nˆÿ/C75/C114T: The quantity /C64T=/C64nis called the gradient of Twhich in vector form is /C114T. We are now ready to derive the equation for heat conduction. Let /C86be an arbitrary volume lying within the solid and bounded by surface S. The total 179HEAT CONDUCTION n Figure 4.20. Heat flux through a thin sheet. amount of heat entering Sper unit time is ZZ S/C75/C114T…†  ^ndS; where ^nis an outward unit vector normal to element surface area dS. Using the divergence theorem, this can be written as ZZ S/C75/C114T…†  ^ndSˆZZZ /C86/C114/C75/C114T…† d/C86: …4:38† Now the heat contained in /C86is given by ZZZ /C86c/C26Td/C86; where cand/C26are respectively the specific heat capacity and density of the solid. Then the time rate of increase of heat is /C64 /C64tZZZ /C86c/C26Td/C86ˆZZZ /C86c/C26/C64T /C64td/C86: …4:39† Equating the right hand sides of Eqs. (4.38) and (4.39) yields ZZZ /C86c/C26/C64T /C64tÿ/C114 /C75/C114T…† d/C86ˆ0: Since /C86is arbitrary, the integrand (assumed continuous) must be identically zero: c/C26/C64T /C64tˆ/C114 /C75/C114T…† or if /C75,c,/C26are constants /C64T /C64tˆk/C114/C114 Tˆk/C1142T; …4:40† where kˆ/C75=c/C26. This is the required equation for heat conduction and was first developed by Fourier in 1822. For the semiinfinite thin bar, the boundary condi- tions are T…x;0†ˆf…x†;T…0;t†ˆ0; jT…x;t†j<M; …4:41† where the last condition means that the temperature must be bounded for physicalreasons. A solution of Eq. (4.40) can be obtained by separation of variables, that is by letting TˆX…x†H…t†: Then XH 0ˆkX00H or X00=XˆH0=kH: 180FOURIER SERIES AND INTEGRALS Each side must be a constant which we call ÿ2. (If we use ‡2, the resulting solution does not satisfy the boundedness condition for real values of .) Then X00‡2Xˆ0;H0‡2kHˆ0 with the solutions X…x†ˆA1cosx‡B1sinx;H…t†ˆC1eÿk2t: A solution to Eq. (4.40) is thus given by T…x;t†ˆC1eÿk2t…A1cosx‡B1sinx† ˆeÿk2t…Acosx‡Bsinx†: From the second of the boundary conditions (4.41) we find Aˆ0 and so T…x;t† reduces to T…x;t†ˆBeÿk2tsinx: Since there is no restriction on the value of , we can replace Bby a function B…† and integrate over from 0 to 1and still have a solution: T…x;t†ˆZ1 0B…†eÿk2tsinxd: …4:42† Using the first of boundary conditions (4.41) we find f…x†ˆZ1 0B…†sinxd: Then by the Fourier sine transform we find B…†ˆ2 Z1 0f…x†sinxdxˆ2 Z1 0f…u†sinudu and the temperature distribution along the semiinfinite thin bar is T…x;t†ˆ2 Z1 0Z1 0f…u†eÿk2tsinusinxddu: …4:43† Using the relation sinusinxˆ1 2‰cos…uÿx†ÿcos…u‡x†Š; Eq. (4.43) can be rewritten T…x;t†ˆ1 Z1 0Z1 0f…u†eÿk2t‰cos…uÿx†ÿcos…u‡x†Šddu ˆ1 Z1 0f…u†Z1 0eÿk2tcos…uÿx†dÿZ1 0eÿk2tcos…u‡x†d du: 181HEAT CONDUCTION Using the integral Z1 0eÿ 2cos/C12dˆ1 2  /C114 eÿ/C122=4 ; we find T…x;t†ˆ1 2 ktpZ1 0f…u†eÿ…uÿx†2=4ktduÿZ1 0f…u†eÿ…u‡x†2=4ktdu : Letting …uÿx†=2ktp ˆ/C119in the first integral and …u‡x†=2 ktp ˆ/C119in the second integral, we obtain T…x;t†ˆ 1pZ1 ÿx=2 ktpeÿ/C1192f…2/C119 ktp ‡x†d/C119ÿZ1 x=2 ktpeÿ/C1192f…2/C119 ktp ÿx†d/C119"# : Fourier transforms for functions of se/C118eral /C118ariables We can extend the development of Fourier transforms to a function of several variables, such as f…x;y;z†. If we first decompose the function into a Fourier integral with respect to x, we obtain f…x;y;z†ˆ 1  2pZ1 ÿ1/C13…/C33x;y;z†ei/C33xxd/C33x; where /C13is the Fourier transform. Similarly, we can decompose the function with respect to yandzto obtain f…x;y;z†ˆ1 …2†2=3Z1 ÿ1/C103…/C33x;/C33y;/C33z†ei…/C33xx‡/C33yy‡/C33zz†d/C33xd/C33yd/C33z; with /C103…/C33x;/C33y;/C33z†ˆ1 …2†2=3Z1 ÿ1f…x;y;z†eÿi…/C33xx‡/C33yy‡/C33zz†dxdydz : We can regard /C33x;/C33y;/C33zas the components of a vector /C33whose magnitude is /C33ˆ /C332x‡/C332y‡/C332zq ; then we express the above results in terms of the vector /C33: f…r†ˆ1 …2†2=3Z1 ÿ1/C103…x†eixrdx; …4:44† /C103…x†ˆ1 …2†2=3Z1 ÿ1f…r†eÿ…ixr†dr: …4:45† 182FOURIER SERIES AND INTEGRALS /C84he Fourier integral and the delta function The delta function is a very useful tool in physics, but it is not a function in the usual mathematical sense. The need for this strange ‘function’ arises naturally from the Fourier integrals. Let us go back to Eqs. (4.30) and (4.31) and substitute /C103…/C33†intof…x†; we then have f…x†ˆ1 2Z1 ÿ1d/C33Z1 ÿ1dx0f…x0†ei/C33…xÿx0†: Interchanging the order of integration gives f…x†ˆZ1 ÿ1dx0f…x0†1 2Z1 ÿ1d/C33ei/C33…xÿx0†: …4:46† If the above equation holds for any function f…x†, then this tells us something remarkable about the integral 1 2Z1 ÿ1d/C33ei/C33…xÿx0† considered as a function of x0. It vanishes everywhere except at x0ˆx, and its integral with respect to x0over any interval including xis unity. That is, we may think of this function as having an infinitely high, infinitely narrow peak atxˆx 0. Such a strange function is called Dirac’s delta function (first introduced by Paul A. M. Dirac): …xÿx0†ˆ1 2Z1 ÿ1d/C33ei/C33…xÿx0†: …4:47† Equation (4.46) then becomes f…x†ˆZ1 ÿ1f…x0†…xÿx0†dx0: …4:48† Equation (4.47) is an integral representation of the delta function. We summarize its properties below: …xÿx0†ˆ0;ifx06ˆx; …4:49a† Zb a…xÿx0†dx0ˆ0;ifx/C62borx<a 1;ifa<x<b;/C26 …4:49b† f…x†ˆZ1 ÿ1f…x0†…xÿx0†dx0: …4:49c† 183THE FOURIER INTEGRAL AND THE DELTA FUNCTION It is often convenient to place the origin at the singular point, in which case the delta function may be written as …x†ˆ1 2Z1 ÿ1d/C33ei/C33x: …4:50† To examine the behavior of the function for both small and large x, we use an alternative representation of this function obtained by integrating as follows: …x†ˆ1 2lim a!1Za ÿaei/C33xd/C33ˆlim a!11 2eiaxÿeÿiax ix ˆlim a!1sinax x; …4:51† where ais positive and real. We see immediately that …ÿx†ˆ…x†. To examine its behavior for small x, we consider the limit as xgoes to zero: lim x!0sinax xˆa lim x!0sinax axˆa : Thus, …0†ˆlima!1…a=†!1 , or the amplitude becomes infinite at the singu- larity. For large jxj, we see that sin …ax†=xoscillates with period 2 =a, and its amplitude falls o/C128 as 1 =jxj. But in the limit as agoes to infinity, the period becomes infinitesimally narrow so that the function approaches zero everywhere except for the infinite spike of infinitesimal width at the singularity. What is the integral of Eq. (4.51) over all space/C63 Z1 ÿ1lim a!1sinax xdxˆlim a!12 Z1 0sinax xdxˆ2  2ˆ1: Thus, the delta function may be thought of as a spike function which has unit area but a non-zero amplitude at the point of singularity, where the amplitude becomes infinite. No ordinary mathematical function with these properties exists. How do we end up with such an improper function/C63 It occurs because the change of order of integration in Eq. (4.46) is not permissible. In spite of this, the Dirac deltafunction is a most convenient function to use symbolically. For in applications the delta function always occurs under an integral sign. Carrying out this integration, using the formal properties of the delta function, is really equivalent to inverting the order of integration once more, thus getting back to a mathematically correct expression. Thus, using Eq. (4.49) we have Z 1 ÿ1f…x†…xÿx0†dxˆf…x0†; but, on substituting Eq. (4.47) for the delta function, the integral on the left hand side becomes Z1 ÿ1f…x†1  2pZ1 ÿ1d/C33ei/C33…xÿx0†/C26/C27 dx 184FOURIER SERIES AND INTEGRALS or, using the property …ÿx†ˆ…x†, Z1 ÿ1f…x†1  2pZ1 ÿ1d/C33eÿi/C33…xÿx0†/C26/C27 dx andchanging the order of integration , we have Z1 ÿ1f…x†1 2pZ 1 ÿ1d/C33eÿi/C33x/C26/C27 ei/C33x0dx: Comparing this expression with Eqs. (4.30) and (4.31), we see at once that this double integral is equal to f…x0†, the correct mathematical expression. It is important to keep in mind that the delta function cannot be the end result of a calculation and has meaning only so long as a subse/C113uent integration over its argu/C45 ment is carried out . We can easily verify the following most frequently required properties of the delta function: Ifa<b Zb af…x†…xÿx0†dxˆf…x0†;ifa<x0<b 0; ifx0<aorx0<b/C26 ; …4:52a† …ÿx†ˆ…x†; …4:52b† 0…x†ˆÿ 0…ÿx†;0…x†ˆd…x†=dx; …4:52c† x…x†ˆ0; …4:52d† …ax†ˆaÿ1…x†;a/C620; …4:52e† …x2ÿa2†ˆ… 2a†ÿ1…xÿa†‡…x‡a† ‰Š ;a/C620; …4:52f† Z …aÿx†…xÿb†dxˆ…aÿb†; …4:52g† f…x†…xÿa†ˆf…a†…xÿa†: …4:52h† Each of the first six of these listed properties can be established by multiplying both sides by a continuous, di/C128erentiable function f…x†and then integrating over x. For example, multiplying x0…x†byf…x†and integrating over xgives Z f…x†x0…x†dxˆÿZ …x†d dx‰xf…x†Šdx ˆÿZ …x†f…x†‡xf0…x†/C2/C3 dxˆÿZ f…x†…x†dx: Thus x…x†has the same e/C128ect when it is a factor in an integrand as has ÿ…x†. 185THE FOURIER INTEGRAL AND THE DELTA FUNCTION Parse/C118al/C39s identit/C121 for Fourier integrals We arrived earlier at Parseval’s identity for Fourier series. An analogy exists for Fourier integrals. If /C103… †and/C71… †are Fourier transforms of f…x†and/C70…x† respectively, we can show that Z1 ÿ1f…x†/C70/C42…x†dxˆ1 2Z1 ÿ1/C103… †/C71/C42… †d ; …4:54† where /C70/C42…x†is the complex conjugate of /C70…x†. In particular, if /C70…x†ˆf…x†and hence /C71… †ˆ/C103… †, then we have Z1 ÿ1f…x†jj2dxˆZ1 ÿ1/C103… †jj d : …4:54† Equation (4.53), or the more general Eq. (4.54), is known as the Parseval’s iden-tity for Fourier integrals. Its proof is straightforward: Z 1 ÿ1f…x†/C70/C42…x†dxˆZ1 ÿ11  2pZ1 ÿ1/C103… †eÿi xd  1 2pZ 1 ÿ1/C71/C42… 0†ei 0xd 0 dx ˆZ1 ÿ1d Z1 ÿ1d 0/C103… †/C71/C42… 0†1 2Z1 ÿ1eix… ÿ 0†dx ˆZ1 ÿ1d /C103… †Z1 ÿ1d 0/C71/C42… 0†… 0ÿ †ˆZ1 ÿ1/C103… †/C71/C42… †d : Parseval’s identity is very useful in understanding the physical interpretation of the transform function /C103… †when the physical significance of f…x†is known. The following example will show this. Example 4.10 Consider the following function, as shown in Fig. 4.21, which might represent thecurrent in an antenna, or the electric field in a radiated wave, or displacement of a damped harmonic oscillator: f…t†ˆ0 t<0 e ÿt=Tsin/C330tt /C620:/C26 186FOURIER SERIES AND INTEGRALS Its Fourier transform /C103…/C33†is /C103…/C33†ˆ1  2pZ1 ÿ1f…t†eÿi/C33tdt ˆ1 2pZ 1 ÿ1eÿt=Teÿi/C33tsin/C330tdt ˆ1 2 2p 1 /C33‡/C330ÿi=Tÿ1 /C33ÿ/C330ÿi=T : Iff…t†is a radiated electric field, the radiated power is proportional to jf…t†j2 and the total energy radiated is proportional toR1 0f…t†jj2dt. This is equal toR1 0/C103…/C33†jj2d/C33by Parseval’s identity. Then j/C103…/C33†j2must be the energy radiated per unit frequency interval. Parseval’s identity can be used to evaluate some definite integrals. As an exam- ple, let us revisit Example 4.8, where the given function is f…x†ˆ1xjj<a 0xjj/C62a( and its Fourier transform is /C103…/C33†ˆ 2 /C114 sin/C33a /C33: By Parseval’s identity, we have Z1 ÿ1f…x†fg2dxˆZ1 ÿ1/C103…/C33†fg2d/C33: 187PARSEVAL’S IDENTITY FOR FOURIER INTEGRALS Figure 4.21. A damped sine wave. This is equivalent to Za ÿa…1†2dxˆZ1 ÿ12 sin2/C33a /C332d/C33; from which we find Z1 02 sin2/C33a /C332d/C33ˆa 2: /C84he con/C118olution theorem for Fourier transforms The convolution of the functions f…x†andH…x†, denoted by f/C3H, is defined by f/C3HˆZ1 ÿ1f…u†H…xÿu†du: …4:55† If/C103…/C33†and/C71…/C33†are Fourier transforms of f…x†andH…x†respectively, we can show that 1 2Z1 ÿ1/C103…/C33†/C71…/C33†ei/C33xd/C33ˆZ1 ÿ1f…u†H…xÿu†du: …4:56† This is known as the convolution theorem for Fourier transforms. It means that the Fourier transform of the product /C103…/C33†/C71…/C33†, the left hand side of Eq. (55), is the convolution of the original function. The proof is not dicult. We have, by definition of the Fourier transform, /C103…/C33†ˆ1  2pZ1 ÿ1f…x†eÿi/C33xdx; /C71…/C33†ˆ1 2pZ 1 ÿ1H…x0†eÿi/C33x0dx0: Then /C103…/C33†/C71…/C33†ˆ1 2Z1 ÿ1Z1 ÿ1f…x†H…x0†eÿi/C33…x‡x0†dxdx0: …4:57† Letx‡x0ˆuin the double integral of Eq. (4.57) and we wish to transform from (x,x0)t o( x;u). We thus have dxdx0ˆ/C64…x;x0† /C64…x;u†dudx ; where the Jacobian of the transformation is /C64…x;x0† /C64…x;u†ˆ/C64x /C64x/C64x /C64u /C64x0 /C64x/C64x0 /C64u/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ10 01/C12/C12/C12/C12/C12/C12/C12/C12ˆ1: 188FOURIER SERIES AND INTEGRALS Thus Eq. (4.57) becomes /C103…/C33†/C71…/C33†ˆ1 2Z1 ÿ1Z1 ÿ1f…x†H…uÿx†eÿi/C33udxdu ˆ1 2Z1 ÿ1eÿi/C33uZ1 ÿ1f…x†H…uÿx†du dx ˆ/C70Z1 ÿ1f…x†H…uÿx†du/C26/C27 ˆ/C70f/C3Hfg : …4:58† From this we have equivalently f/C3Hˆ/C70ÿ1/C103…/C33†/C71…/C33† fg ˆ … 1=2†Z1 ÿ1ei/C33x/C103…/C33†/C71…/C33†; which is Eq. (4.56). Equation (4.58) can be rewritten as /C70ffg/C70Hfgˆ/C70f/C3Hfg …/C103ˆ/C70ffg;/C71ˆ/C70Hfg † ; which states that the /C70ourier transform of the convolution of f(x) and /C72(x) is e/C113ual to the product of the /C70ourier transforms of f(x) and /C72(x) . This statement is often taken as the convolution theorem. The convolution obeys the commutative, associative and distributive laws of algebra that is, if we have functions f1;f2;f3then f1/C3f2ˆf2/C3f1 commutative; f1/C3…f2/C3f3†ˆ… f1/C3f2†/C3f3 associative; f1/C3…f2‡f3†ˆf1/C3f2‡f1/C3f3distributive :9 >>= >>;…4:59† It is not dicult to prove these relations. For example, to prove the commutative law, we first have f1/C3f2Z1 ÿ1f1…u†f2…xÿu†du: Now let xÿuˆ/C118, then f1/C3f2Z1 ÿ1f1…u†f2…xÿu†du ˆZ1 ÿ1f1…xÿ/C118†f2…/C118†d/C118ˆf2/C3f1: Example 4.11Solve the integral equation y…x†ˆf…x†‡R 1 ÿ1y…u†r…xÿu†du, where f…x†and r…x†are given, and the Fourier transforms of y…x†;f…x†andr…x†exist. 189THE CONVOLUTION THEOREM FOR FOURIER TRANSFORMS Solution: Let us denote the Fourier transforms of y…x†;f…x†and r…x†by /C89…/C33†;/C70…/C33†;and /C82…/C33†respectively. Taking the Fourier transform of both sides of the given integral equation, we have by the convolution theorem Y…/C33†ˆ/C70…/C33†‡Y…/C33†R…/C33†orY…/C33†ˆ/C70…/C33† 1ÿR…/C33†: /C67alculations of Fourier transforms Fourier transforms can often be used to transform a di/C128erential equation which is dicult to solve into a simpler equation that can be solved relatively easy. In order to use the transform methods to solve first- and second-order di/C128erential equations, the transforms of first- and second-order derivatives are needed. By taking the Fourier transform with respect to the variable x, we can show that …a†/C70/C64u /C64x ˆi /C70…u†; …b†/C70/C642u /C64x2/C32! ˆÿ 2/C70…u†; …c†/C70/C64u /C64t ˆ/C64 /C64t/C70…u†:9 >>>>>>>>>= >>>>>>>>>;…4:60† Proof: (a) By definition we have /C70 /C64u /C64x ˆZ1 ÿ1/C64u /C64xeÿi xdx; where the factor 1 =  2p has been dropped. Using integration by parts, we obtain /C70/C64u /C64x ˆZ1 ÿ1/C64u /C64xeÿi xdx ˆueÿi x/C12/C12/C12/C121 ÿ1‡i Z1 ÿ1ueÿi xdx ˆi /C70…u†: (b) Let uˆ/C64/C118=/C64xin (a), then /C70/C642/C118 /C64x2/C32! ˆi /C70/C64/C118 /C64x ˆ…i †2/C70…/C118†: Now if we formally replace /C118byuwe have /C70/C642u /C64x2/C32! ˆÿ 2/C70…u†; 190FOURIER SERIES AND INTEGRALS provided that uand/C64u=/C64x!0a sx! 1 . In general, we can show that /C70/C64nu /C64xn ˆ…i †n/C70…u† ifu;/C64u=/C64x;...;/C64nÿ1u=/C64xnÿ1! 1 asx! 1 . (c) By definition /C70/C64u /C64t ˆZ1 ÿ1/C64u /C64teÿi xdxˆ/C64 /C64tZ1 ÿ1ueÿi xdxˆ/C64 /C64t/C70…u†: Example 4.12 Solve the inhomogeneous di/C128erential equation d2 dx2‡/C112d dx‡/C113/C32! f…x†ˆR…x†;ÿ1 x1 ; where pand/C113are constants. Solution: We transform both sides /C70d2f dx2‡/C112df dx‡/C113f() ˆ‰ …i †2‡/C112…i †‡/C113Š/C70f…x†fg ˆ/C70R…x†fg : If we denote the Fourier transforms of f…x†andR…x†by/C103… †and/C71… †, respec- tively, /C70f…x†fg ˆ/C103… †;/C70R…x†fg ˆ/C71… †; we have …ÿ 2‡i/C112 ‡/C113†/C103… †ˆ/C71… †;or/C103… †ˆ/C71… †=…ÿ 2‡i/C112 ‡/C113† and hence f…x†ˆ1  2pZ1 ÿ1ei x/C103… †d ˆ1 2pZ 1 ÿ1ei x /C71… † ÿ 2‡i/C112 ‡/C113d : We will not gain anything if we do not know how to evaluate this complex integral. This is not a dicult problem in the theory of functions of complex variables (see Chapter 7). 191CALCULATIONS OF FOURIER TRANSFORMS /C84he delta function and the /C71reen/C39s function method The Green’s function method is a very useful technique in the solution of partial di/C128erential equations. It is usually used when boundary conditions, rather than initial conditions, are specified. To appreciate its usefulness, let us consider the inhomogeneous di/C128erential equation L…x†f…x†ÿf…x†ˆR…x†… 4:61† over a domain /C68, with Lan arbitrary di/C128erential operator, and a given con- stant. Suppose we can expand f…x†andR…x†in eigenfunctions unof the operator L…Lunˆnun†: f…x†ˆX ncnun…x†;R…x†ˆX ndnun…x†: Substituting these into Eq. (4.61) we obtain X ncn…nÿ†un…x†ˆX ndnun…x†: Since the eigenfunctions un…x†are linearly independent, we must have cn…nÿ†ˆdnorcnˆdn=…nÿ†: Moreover, dnˆZ Dun/C42R…x†dx: Now we may write cnas cnˆ1 nÿZ Dun/C42R…x†dx; therefore f…x†ˆX nun nÿZ Dun/C42…x0†R…x0†dx0: This expression may be written in the form f…x†ˆZ D/C71…x;x0†R…x0†dx0; …4:62† where /C71…x;x0†is given by /C71…x;x0†ˆX nun…x†un/C42…x0† nÿ…4:63† and is called the Green’s function. Some authors prefer to write /C71…x;x0;†to emphasize the dependence of /C71onas well as on xandx0. 192FOURIER SERIES AND INTEGRALS What is the di/C128erential equation obeyed by /C71…x;x0†/C63 Suppose f…x0†in Eq. (4.62) is taken to be …x0ÿx0), then we obtain f…x†ˆZ D/C71…x;x0†…x0ÿx0†dxˆ/C71…x;x0†: Therefore /C71…x;x0†is the solution of L/C71…x;x0†ÿ/C71…x;x0†ˆ…xÿx0†; …4:64† subject to the appropriate boundary conditions. Eq. (4.64) shows clearly that the /C71reen/C39s function is the solution of the problem for a unit point /C96source/C39 R…x†ˆ…xÿx0†. Example 4.13 Find the solution to the di/C128erential equation d2u dx2ÿk2uˆf…x†… 4:65† on the interval 0 x/C108, with u…0†ˆu…/C108†ˆ0, for a general function f…x†. Solution: We first solve the di/C128erential equation which /C71…x;x0†obeys: d2/C71…x;x0† dx2ÿk2/C71…x;x0†ˆ…xÿx0†: …4:66† Forxequal to anything but x0(that is, for x<x0orx/C62x0),…xÿx0†ˆ0a n d we have d2/C71<…x;x0† dx2ÿk2/C71<…x;x0†ˆ0 …x<x0†; d2/C71/C62…x;x0† dx2ÿk2/C71/C62…x;x0†ˆ0 …x/C62x0†: Therefore, for x<x0 /C71<ˆAekx‡Beÿkx: By the boundary condition u…0†ˆ0 we find A‡Bˆ0, and /C71<reduces to /C71<ˆA…ekxÿeÿkx†; …4:67a† similarly, for x/C62x0 /C71/C62ˆCekx‡Deÿkx: 193THE DELTA FUNCTION AND THE GREEN’S FUNCTION METHOD By the boundary condition u…/C108†ˆ0 we find Cek/C108‡Deÿk/C108ˆ0, and /C71/C62can be rewritten as /C71/C62ˆC0‰ek…xÿ/C108†ÿeÿk…xÿ/C108†Š; …4:67b† where C0ˆCek/C108. How do we determine the constants AandC0/C63 First, continuity of /C71atxˆx0 gives A…ekxÿeÿkx†ˆC0…ek…xÿ/C108†ÿeÿk…xÿ/C108††: …4:68† A second constraint is obtained by integrating Eq. (4.61) from x0ÿ/C34tox0‡/C34, where /C34is infinitesimal: Zx0‡/C34 x0ÿ/C34d2/C71 dx2ÿk2/C71"# dxˆZx0‡/C34 x0ÿ/C34…xÿx0†dxˆ1: …4:69† But Zx0‡/C34 x0ÿ/C34k2/C71dxˆk2…/C71/C62ÿ/C71<†ˆ0; where the last step is required by the continuity of /C71. Accordingly, Eq. (4.64) reduces to Zx0‡/C34 x0ÿ/C34d2/C71 dx2dxˆd/C71/C62 dxÿd/C71< dxˆ1: …4:70† Now d/C71< dx xˆx0/C12/C12/C12/C12/C12ˆAk…ekx0‡eÿkx0† and d/C71/C62 dx xˆx0/C12/C12/C12/C12/C12ˆC 0k‰ek…x0ÿ/C108†‡eÿk…x0ÿ/C108†Š: Substituting these into Eq. (4.70) yields C0k…ek…x0ÿ/C108†‡eÿk…x0ÿ/C108††ÿAk…ekx0‡eÿkx0†ˆ1: …4:71† We can solve Eqs. (4.68) and (4.71) for the constants AandC0. After some algebraic manipulation, the solution is Aˆ1 2ksinhk…x0ÿ/C108† sinhk/C108;C0ˆ1 2ksinhkx0 sinhk/C108 194FOURIER SERIES AND INTEGRALS and the Green’s function is /C71…x;x0†ˆ1 ksinhkxsinhk…x0ÿ/C108† sinhk/C108; …4:72† which can be combined with f…x†to obtain u…x†: u…x†ˆZ/C108 0/C71…x;x0†f…x0†dx0: Problems 4.1 ( a) Find the period of the function f…x†ˆcos…x=3†‡cos…x=4†. (b) Show that, if the function f…t†ˆcos/C331t‡cos/C332tis periodic with a period T, then the ratio /C331=/C332must be a rational number. 4.2 Show that if f…x‡P†ˆf…x†, then Za‡P=2 aÿP=2f…x†dxˆZP=2 ÿP=2f…x†dx;ZP‡x Pf…x†dxˆZx 0f…x†dx: 4.3 ( a) Using the result of Example 4.2, prove that 1ÿ1 3‡15ÿ17‡ÿˆ 4: (b) Using the result of Example 4.3, prove that 1 13ÿ1 35‡1 57ÿ‡ˆÿ2 4: 4.4 Find the Fourier series which represents the function f…x†ˆjxjin the inter- valÿx. 4.5 Find the Fourier series which represents the function f…x†ˆxin the interval ÿx. 4.6 Find the Fourier series which represents the function f…x†ˆx2in the inter- valÿx. 4.7 Represent f…x†ˆx;0<x<2, as: ( a) in a half-range sine series, ( b) a half- range cosine series. 4.8 Represent f…x†ˆsinx,0<x<, as a Fourier cosine series. 4.9 ( a) Show that the function f…x†of period 2 which is equal to xon…ÿ1;1† can be represented by the following Fourier series ÿi eixÿeÿixÿ12e 2ix‡12e ÿ2ix‡13e 3ixÿ13e ÿ3ix‡ : (b) Write Parseval’s identity corresponding to the Fourier series of ( a). (c) Determine from ( b) the sum Sof the series 1 ‡1 4‡19‡ˆP1 nˆ11=n2. 195PROBLEMS 4.10 Find the exponential form of the Fourier series of the function whose defi- nition in one period is f…x†ˆeÿx;ÿ1<x<1. 4.11 ( a) Show that the set of functions 1;sinx L;cosx L;sin2x L;cos2x L;sin3x L;cos3x L;... form an orthogonal set in the interval …ÿL;L†. (b) Determine the corresponding normalizing constants for the set in ( a)s o that the set is orthonormal in …ÿL;L†. 4.12 Express f…x;y†ˆxyas a Fourier series for 0 x1;0y2. 4.13 Steady-state heat conduction in a rectangular plate: Consider steady-state heat conduction in a flat plate having temperature values prescribed on the sides (Fig. 4.22). The boundary value problem modeling this is: /C642u /C642x2‡/C642u /C642y2ˆ0; 0<x< ; 0<y</C12 ; u…x;0†ˆu…x;/C12†ˆ0; 0<x< ; u…0;y†ˆ0;u… ;y†ˆT; 0<y</C12 : Determine the temperature at any point of the plate. 4.14 Derive and solve the following eigenvalue problem which occurs in the theory of a vibrating square membrane whose sides, of length L, are kept fixed: /C642/C119 /C64x2‡/C642/C119 /C64y2‡/C119ˆ0; /C119…0;y†ˆ/C119…L;y†ˆ0…0yL†; /C119…x;0†ˆ/C119…x;L†ˆ0…0yL†: 196FOURIER SERIES AND INTEGRALS Figure 4.22. Flat plate with prescribed temperature. 4.15 Show that the Fourier integral can be written in the form f…x†ˆ1 Z1 0d/C33Z1 ÿ1f…x0†cos/C33…xÿx0†dx0: 4.16 Starting with the form obtained in Problem 4.15, show that the Fourier integral can be written in the form f…x†ˆZ1 0A…/C33†cos/C33x‡B…/C33†sin/C33x fg d/C33; where A…/C33†ˆ1 Z1 ÿ1f…x†cos/C33xd x; B…/C33†ˆ1 Z1 ÿ1f…x†sin/C33xd x: 4.17 ( a) Find the Fourier transform of f…x†ˆ1ÿx2jxj<1 0 jxj/C621:( (b) Evaluate Z1 0xcosxÿsinx x3cosx 2dx: 4.18 ( a) Find the Fourier cosine transform of f…x†ˆeÿmx;m/C620. (b) Use the result in ( a) to show that Z1 0cos/C112x x2‡ 2dxˆ 2 eÿ/C112 …/C112/C620; /C62 0†: 4.19 Solve the integral equation Z1 0f…x†sin xd xˆ1ÿ 0 1 0 /C621/C26 : 4.20 Find a bounded solution to Laplace’s equation /C1142u…x;y†ˆ0 for the half- plane y/C620i futakes on the value of f(x) on the x-axis: /C642u /C64x2‡/C642u /C64y2ˆ0; u…x;0†ˆf…x†; u…x;y† jj <M: 4.21 Show that the following two functions are valid representations of the delta function, where /C34is positive and real: …a†…x†ˆ1plim /C34!01/C34peÿx2=/C34 …b†…x†ˆ1 lim /C34!0/C34 x2‡/C342: 197PROBLEMS 4.22 Verify the following properties of the delta function: (a)…x†ˆ…ÿx†, (b)x…x†ˆ0, (c)0…ÿx†ˆÿ 0…x†, (d)x0…x†ˆÿ …x†, (e)c…cx†ˆ…x†;c/C620. 4.23 Solve the integral equation for y…x† Z1 ÿ1y…u†du …xÿu†2‡a2ˆ1 x2‡b20<a<b: 4.24 Use Fourier transforms to solve the boundary value problem /C64u /C64tˆk/C642u /C64x2; u…x;0†ˆf…x†; u…x;t† jj <M; where ÿ1 <x<1;t/C620. 4.25 Obtain a solution to the equation of a driven harmonic oscillator /C127x…t†‡2/C12_x…t†‡/C332 0x…t0ˆR…t†; where /C12and/C330are positive and real constants. 198FOURIER SERIES AND INTEGRALS 5 Linear vector spaces Linear vector space is to quantum mechanics what calculus is to classical mechanics. In this chapter the essential ideas of linear vector spaces will be dis- cussed. The reader is already familiar with vector calculus in three-dimensional Euclidean space /C693(Chapter 1). We therefore present our discussion as a general- ization of elementary vector calculus. The presentation will be, however, slightlyabstract and more formal than the discussion of vectors in Chapter 1. Any reader who is not already familiar with this sort of discussion should be patient with the first few sections. You will then be amply repaid by finding the rest of this chapter relatively easy reading. /C69uclidean n-space E n In the study of vector analysis in /C693, an ordered triple of numbers ( a1,a2,a3) has two di/C128erent geometric interpretations. It represents a point in space, with a1,a2, a3being its coordinates; it also represents a vector, with a1,a2, and a3being its components along the three coordinate axes (Fig. 5.1). This idea of using triples of numbers to locate points in three-dimensional space was first introduced in the mid-seventeenth century. By the latter part of the nineteenth century physicists and mathematicians began to use the quadruples of numbers ( a1,a2,a3,a4)a s points in four-dimensional space, quintuples ( a1,a2,a3,a4,a5) as points in five- dimensional space etc. We now extend this to n-dimensional space /C69n, where nis a positive integer. Although our geometric visualization doesn’t extend beyond three-dimensional space, we can extend many familiar ideas beyond three-dimen- sional space by working with analytic or numerical properties of points and vectors rather than their geometric properties. For two- or three-dimensional space, we use the terms ‘ordered pair’ and ‘ordered triple.’ When n/C623, we use the term ‘ordered- n-tuplet’ for a sequence 199 ofnnumbers, real or complex, ( a1,a2,a3;...;an); they will be viewed either as a generalized point or a generalized vector in a n-dimensional space /C69n. Two vectors uˆ…u1;u2;...;un†and/C118ˆ…/C1181;/C1182;...;/C118n†in/C69nare called equal if uiˆ/C118i;iˆ1;2;...;n …5:1† The sum u‡/C118is defined by u‡/C118ˆ…u1‡/C1181;u2‡/C1182;...;un‡/C118n†… 5:2† and if kis any scalar, the scalar multiple kuis defined by kuˆ…ku1;ku2;...;kun†: …5:3† Ifuˆ…u1;u2;...;un†is any vector in /C69n, its negative is given by ÿuˆ… ÿ u1;ÿu2;...;ÿun†… 5:4† and the subtraction of vectors in /C69ncan be considered as addition: /C118ÿuˆ /C118‡… ÿ u†. The null (zero) vector in /C69nis defined to be the vector 0 ˆ…0;0;...;0†. The addition and scalar multiplication of vectors in /C69nhave the following arithmetic properties: u‡/C118ˆ/C118‡u; …5:5a† u‡…/C118‡/C119†ˆ… u‡/C118†‡/C119; …5:5b† u‡/C48ˆ/C48‡uˆu; …5:5c† a…bu†ˆ…ab†u; …5:5d† a…u‡/C118†ˆau‡a/C118; …5:5e† …a‡b†uˆau‡bu; …5:5f† where u,/C118,/C119are vectors in /C69nand aand bare scalars. 200LINEAR VECTOR SPACES Figure 5.1. A space point Pwhose position vector is /C65. We usually define the inner product of two vectors in /C693in terms of lengths of the vectors and the angle between the vectors: /C65/C66ˆABcos; ˆ/C128… /C65;/C66†.W e do not define the inner product in /C69nin the same manner. However, the inner product in /C693has a second equivalent expression in terms of components: /C65/C66ˆA1B1‡A2B2‡A3B3. We choose to define a similar formula for the gen- eral case. We made this choice because of the further generalization that will be outlined in the next section. Thus, for any two vectors uˆ…u1;u2;...;un†and /C118ˆ…/C1181;/C1182;...;/C118n†in/C69n, the inner (or dot) product u/C118is defined by u/C118ˆu1/C42/C1181‡u2/C42/C1182‡‡ un/C42/C118n …5:6† where the asterisk denotes complex conjugation. uis often called the prefactor and/C118the post-factor. The inner product is linear with respect to the post-factor, and anti-linear with respect to the prefactor: u…a/C118‡b/C119†ˆau/C118‡bu/C119;…au‡b/C118†/C119ˆa/C42…u/C118†‡b/C42…u/C119†: We expect the inner product for the general case also to have the following three main features: u/C118ˆ…/C118u†/C42 …5:7a† u…a/C118‡b/C119†ˆau/C118‡bu/C119 …5:7b† uu0…ˆ0;if and only if uˆ0†: …5:7c† Many of the familiar ideas from /C692and/C693have been carried over, so it is common to refer to /C69nwith the operations of addition, scalar multiplication, and with the inner product that we have defined here as Euclidean n-space. /C71eneral linear /C118ector spaces We now generalize the concept of vector space still further: a set of ‘objects’ (or elements) obeying a set of axioms, which will be chosen by abstracting the most important properties of vectors in /C69n, forms a linear vector space /C86nwith the objects called vectors. Before introducing the requisite axioms, we first adapt a notation for our general vectors: general vectors are designated by the symbol ji, which we call, following Dirac, ket vectors; the conjugates of ket vectors aredenoted by the symbol hj, the bra vectors. However, for simplicity, we shall refer in the future to the ket vectors jisimply as vectors, and to the hjsa s conjugate vectors. We now proceed to define two basic operations on these vectors: addition and multiplication by scalars. By addition we mean a rule for forming the sum, denoted j/C32 1i‡j/C322i, for any pair of vectors j/C321iandj/C322i. By scalar multiplication we mean a rule for associating with each scalar k and each vector j/C32ia new vector kj/C32i. 201GENERAL LINEAR VECTOR SPACES We now proceed to generalize the concept of a vector space. An arbitrary set of nobjects j1i;j2i;j3i;...;j/C30i;...;j’iform a linear vector /C86nif these objects, called vectors, meet the following axioms or properties: A.1 If /C30jiand ’jiare objects in /C86nandkis a scalar, then /C30ji‡’jiandk/C30jiare in/C86n, a feature called closure. A.2 /C30ji‡’jiˆ’ji‡/C30ji; that is, addition is commutative. A.3 ( /C30ji‡’ji †‡/C32jiˆ/C30ji‡…’ji‡/C32ji); that is, addition is associative. A.4k…/C30ji‡’ji †ˆk/C30ji‡k’ji; that is, scalar multiplication is distributive in the vectors. A.5…k‡ †/C30jiˆk/C30ji‡ /C30ji; that is, scalar multiplication is distributive in the scalars. A.6k… /C30ji †ˆk /C30ji; that is, scalar multiplication is associative. A.7 There exists a null vector 0 jiin/C86nsuch that /C30ji‡0jiˆ/C30jifor all /C30jiin/C86n. A.8 For every vector /C30jiin/C86n, there exists an inverse under addition, ÿ/C30ji such that /C30ji‡ÿ /C30ji ˆ0ji. The set of numbers a;b;...used in scalar multiplication of vectors is called the field over which the vector field is defined. If the field consists of real numbers, we have a real vector field; if they are complex, we have a complex field. /C78ote that the vectors themselves are neither real nor complex/C44 the nature of the vectors is notspeci/C174ed. /C86ectors can be any kinds of objects/C59 all that is re/C113uired is that the vector space axioms be satis/C174ed . Thus we purposely do not use the symbol /C86to denote the vectors as the first step to turn the reader away from the limited concept of the vector as a directed line segment. Instead, we use Dirac’s ket and bra symbols, ji and jh, to denote generic vectors. The familiar three-dimensional space of position vectors /C69 3is an example of a vector space over the field of real numbers. Let us now examine two simpleexamples. Example 5.1 Let/C86be any plane through the origin in /C69 3. We wish to show that the points in the plane /C86form a vector space under the addition and scalar multiplication operations for vector in /C693. Solution: Since E3itself is a vector space under the addition and scalar multi- plication operations, thus Axioms A.2, A.3, A.4, A.5, and A.6 hold for all points inE3and consequently for all points in the plane /C86. We therefore need only show that Axioms A.1, A.7, and A.8 are satisfied. Now the plane /C86, passing through the origin, has an equation of the form ax1‡bx2‡cx3ˆ0: 202LINEAR VECTOR SPACES Hence, if uˆ…u1;u2;u3†and/C118ˆ…/C1181;/C1182;/C1183†are points in /C86, then we have au1‡bu2‡cu3ˆ0a n d a/C1181‡b/C1182‡c/C1183ˆ0: Addition gives a…u1‡/C1181†‡b…u2‡/C1182†‡c…u3‡/C1183†ˆ0; which shows that the point u‡/C118also lies in the plane /C86. This proves that Axiom A.1 is satisfied. Multiplying au1‡bu2‡cu3ˆ0 through by ÿ1 gives a…ÿu1†‡b…ÿu2†‡c…ÿu3†ˆ0; that is, the point ÿuˆ… ÿ u1;ÿu2;ÿu3†lies in /C86. This establishes Axiom A.8. The verification of Axiom A.7 is left as an exercise. Example 5.2 Let/C86be the set of all mnmatrices with real elements. We know how to add matrices and multiply matrices by scalars. The corresponding rules obey closure,associativity and distributive requirements. The null matrix has all zeros in it, and the inverse under matrix addition is the matrix with all elements negated. Thus the set of all mnmatrices, together with the operations of matrix addition and scalar multiplication, is a vector space. We shall denote this vector space by thesymbol M mn. /C83ubspaces Consider a vector space /C86.I f/C87is a subset of /C86and forms a vector space under the addition and scalar multiplication, then /C87is called a subspace of /C86. For example, lines and planes passing through the origin form vector spaces andthey are subspaces of /C69 3. Example 5.3We can show that the set of all 2 2 matrices having zero on the main diagonal is a subspace of the vector space M 22of all 2 2 matrices. Solution: To prove this, let ~Xˆ0x12 x21 0/C32! ~Yˆ0y12 y21 0/C32! be two matrices in /C87and kany scalar. Then k~Xˆ0x12 kx21 0/C32! and ~X‡~Yˆ0 x12‡y12 x21‡y21 0/C32! and thus they lie in /C87. We leave the verification of other axioms as exercises. 203SUBSPACES Linear combination A vector /C87jiis a linear combination of the vectors /C1181ji;/C1182ji;...;/C118rjiif it can be expressed in the form /C87jiˆk1j/C1181i‡k2j/C1182i‡‡ krj/C118ri; where k1;k2;...;krare scalars. For example, it is easy to show that the vector /C87jiˆ…9;2;7†in/C693is a linear combination of /C1181jiˆ…1;2;ÿ1†and /C1182jiˆ…6;4;2†. To see this, let us write …9;2;7†ˆk1…1;2;ÿ1†‡k2…6;4;2† or …9;2;7†ˆ…k1‡6k2;2k1‡4k2;ÿk1‡2k2†: Equating corresponding components gives k1‡6k2ˆ9;2k1‡4k2ˆ2;ÿk1‡2k2ˆ7: Solving this system yields k1ˆÿ3 and k2ˆ2 so that /C87jiˆÿ3/C1181ji‡2/C1182ji: Linear independence/C44 bases/C44 and dimensionalit/C121 Consider a set of vectors 1 ji;2ji;...;rji;...njiin a linear vector space /C86. If every vector in /C86is expressible as a linear combination of 1 ji;2ji;...;rji;...;nji, then we say that these vectors span the vector space /C86, and they are called the base vectors orbasis of the vector space /C86. For example, the three unit vectors e1ˆ…1;0;0†;e2ˆ…0;1;0†, and e3ˆ…0;0;1†span /C693because every vector in /C693 is expressible as a linear combination of e1,e2, and e3. But the following three vectors in /C693do not span /C693/C581jiˆ…1;1;2†;2jiˆ…1;0;1†, and 3 jiˆ…2;1;3†. Base vectors are very useful in a variety of problems since it is often possible to study a vector space by first studying the vectors in a base set, then extending the results to the rest of the vector space. Therefore it is desirable to keep the spanning set as small as possible. Finding the spanning sets for a vector space depends upon the notion of linear independence. We say that a finite set of nvectors 1 ji;2ji;...;rji;...;nji, none of which is a null vector, is linearly independent if no set of non-zero numbers akexists such that Xn kˆ1akkijˆj0i: …5:8† In other words, the set of vectors is linearly independent if it is impossible to construct the null vector from a linear combination of the vectors except when all 204LINEAR VECTOR SPACES the coecients vanish. For example, non-zero vectors 1 jiand 2jiof/C692that lie along the same coordinate axis, say x1, are not linearly independent, since we can write one as a multiple of the other: 1 jiˆa2ji, where ais a scalar which may be positive or negative. That is, 1 jiand 2jidepend on each other and so they are not linearly independent. Now let us move the term a2jito the left hand side and the result is the null vector: 1 jiÿa2jiˆ0ji. Thus, for these two vectors 1 jiand 2jiin /C692, we can find two non-zero numbers (1, ÿa†such that Eq. (5.8) is satisfied, and so they are not linearly independent. On the other hand, the nvectors 1 ji;2ji;...;rji;...;njiare linearly dependent if it is possible to find scalars a1;a2;...;an, at least two of which are non-zero, such that Eq. (5.8) is satisfied. Let us say a96ˆ0. Then we could express 9 jiin terms of the other vectors 9jiˆXn iˆ1;6ˆ9ÿai a9iji: That is, the nvectors in the set are linearly dependent if any one of them can be expressed as a linear combination of the remaining nÿ1 vectors. Example 5.4 The set of three vectors 1 jiˆ…2;ÿ1;0;3†,2jiˆ…1;2;5;ÿ1†;3jiˆ…7;ÿ1;5;8†is linearly dependent, since 3 1 ji ‡ 2ji ÿ 3ji ˆ 0ji. Example 5.5The set of three unit vectors e 1jiˆ…1;0;0†;e2jiˆ…0;1;0†, and e3jiˆ…0;0;1†in /C693is linearly independent. To see this, let us start with Eq. (5.8) which now takes the form a1e1ji‡a2e2ji‡a3e3jiˆ0ji or a1…1;0;0†‡a2…0;1;0†‡a3…0;0;1†ˆ… 0;0;0† from which we obtain …a1;a2;a3†ˆ… 0;0;0†; the set of three unit vectors e1ji;e2ji, and e3jiis therefore linearly independent. Example 5.6 The set Sof the following four matrices 1jiˆ10 00 ;2jiˆ0100 ;j3iˆ0010 ;4jiˆ0001 ; 205LINEAR INDEPENDENCE, BASES, AND DIMENSIONALITY is a basis for the vector space M22of 22 matrices. To see that Sspans M22, note that a typical 2 2 vector (matrix) can be written as ab cd/C32! ˆa10 00/C32! ‡b0100/C32! ‡c0010/C32! ‡d0001/C32! ˆa1ji‡b2ji‡c3ji‡d4ji: To see that Sis linearly independent, assume that a1ji‡b2ji‡c3ji‡d4jiˆ0ji; that is, a10 00 ‡b0100 ‡c0010 ‡d0001 ˆ0000 ; from which we find aˆbˆcˆdˆ0 so that Sis linearly independent. We now come to the dimensionality of a vector space. We think of space around us as three-dimensional. How do we extend the notion of dimension to a linear vector space/C63 Recall that the three-dimensional Euclidean space /C69 3is spanned by the three base vectors: e1ˆ…1;0;0†,e2ˆ…0;1;0†,e3ˆ…0;0;1†. Similarly, the dimension nof a vector space /C86is defined to be the number nof linearly independent base vectors that span the vector space /C86. The vector space will be denoted by /C86n…R†if the field is real and by /C86n…C†if the field is complex. For example, as shown in Example 5.6, 2 2 matrices form a four-dimensional vector space whose base vectors are 1jiˆ10 00 ;2jiˆ0100 ;3jiˆ0010 ;4jiˆ0001 ; since any arbitrary 2 2 matrix can be written in terms of these: ab cd ˆa1ji‡b2ji‡c3ji‡d4ji: If the scalars a;b;c;dare real, we have a real four-dimensional space, if they are complex we have a complex four-dimensional space. Inner product spaces (unitar/C121 spaces) In this section the structure of the vector space will be greatly enriched by the addition of a numerical function, the inner product (or scalar product). Linear vector spaces in which an inner product is defined are called inner-product spaces (or unitary spaces). The study of inner-product spaces enables us to make a real juncture with physics. 206LINEAR VECTOR SPACES In our earlier discussion, the inner product of two vectors in /C69nwas defined by Eq. (5.6), a generalization of the inner product of two vectors in /C693. In a general linear vector space, an inner product is defined axiomatically analogously with the inner product on /C69n. Thus given two vectors Ujiand /C87ji UjiˆXn iˆ1uiiji;/C87jiˆXn iˆ1/C119iiji; …5:9† where Ujiand /C87ji are expressed in terms of the nbase vectors iji, the inner product, denoted by the symbol U/C87jih , is defined to be hU/C87jiˆXn iˆ1Xn jˆ1ui/C42/C119jijjih: …5:10† Uhjis often called the pre-factor and /C87jithe post-factor. The inner product obeys the following rules (or axioms): B.1 U/C87jiˆ/C87Uji/C42 h h (skew-symmetry); B.2 UjUhi 0;ˆ0 if and only if Ujiˆ0ji (positive semidefiniteness); B.3 UXji‡/C87ji † ˆUXji‡U/C87jih h h (additivity); B.4 aU /C87jiˆa/C42U/C87ji;Ub /C87ji ˆbU/C87jih h h h (homogeneity); where aandbare scalars and the asterisk (/C42) denotes complex conjugation. Note that Axiom B.1 is di/C128erent from the one for the inner product on /C693: the inner product on a general linear vector space depends on the order of the two factors for a complex vector space. In a real vector space /C693, the complex conjugation in Axioms B.1 and B.4 adds nothing and may be ignored. In either case, real orcomplex, Axiom B.1 implies that UUjih is real, so the inequality in Axiom B.2 makes sense. The inner product is linear with respect to the post-factor: Uhja/C87‡bXiˆaUhj/C87i‡bUhjXi; and anti-linear with respect to the prefactor, aU‡bX /C87jiˆa/C42U/C87ji‡b/C42X/C87ji: h h h Two vectors are said to be orthogonal if their inner product vanishes. And we will refer to the quantity hUUji 1=2ˆkUkas the norm or length of the vector. A normalized vector, having unit norm, is a unit vector. Any given non-zero vector may be normalized by dividing it by its length. An orthonormal basis is a set ofbasis vectors that are all of unit norm and pair-wise orthogonal. It is very handy to have an orthonormal set of vectors as a basis for a vector space, so for hijjiin Eq. (5.10) we shall assume ihjjiˆ ijˆ1 for iˆj 0 for i6ˆj( ; 207INNER PRODUCT SPACES (UNITARY SPACES) then Eq. (5.10) reduces to Uhj/C87iˆX iX jui/C42/C119jijˆX iui/C42X j/C119jij/C32! ˆX iui/C42/C119i: …5:11† Note that Axiom B.2 implies that if a vector Ujiis orthogonal to every vector of the vector space, then Ujiˆ0: since Uˆ0ijh for all jibelongs to the vector space, so we have in particular UUji ˆ 0 h . We will show shortly that we may construct an orthonormal basis from an arbitrary basis using a technique known as the Gram–Schmidt orthogonalization process. Example 5.7 LetjUiˆ… 3ÿ4i†j1i‡…5ÿ6i†j2iandj/C87iˆ… 1ÿi†j1i‡…2ÿ3i†j2ibe two vec- tors expanded in terms of an orthonormal basis j1iandj2i. Then we have, using Eq. (5.10): UhjUiˆ…3‡4i†…3ÿ4i†‡…5‡6i†…5ÿ6i†ˆ86; /C87hj/C87iˆ…1‡i†…1ÿi†‡…2‡3i†…2ÿ3i†ˆ15; Uhj/C87iˆ…3‡4i†…1ÿi†‡…5‡6i†…2ÿ3i†ˆ35ÿ2iˆ/C87hjUi/C42: Example 5.8If~Aand ~Bare two matrices, where ~Aˆa 11a12 a21a22/C32! ; ~Bˆb11b12 b21b22/C32! ; then the following formula defines an inner product on M22: ~A/C10/C12/C12~B/C11 ˆa11b11‡a12b12‡a21b21‡a22b22: To see this, let us first expand ~Aand ~Bin terms of the following base vectors 1ji ˆ10 00 ;2ji ˆ0100 ;3ji ˆ0010 ;4ji ˆ0001 ; ~Aˆa 111ji‡a122ji‡a213ji‡a224ji; ~Bˆb111ji‡b122ji‡b213ji‡b224ji: The result follows easily from the defining formula (5.10). Example 5.9 Consider the vector jUi, in a certain orthonormal basis, with components Ujiˆ1‡i 3p ‡i/C32! ;iˆ  ÿ1p : 208LINEAR VECTOR SPACES We now expand it in a new orthonormal basis je1i;je2iwith components e1jiˆ1 2p1 1 ;e2jiˆ1 2p1 ÿ1 : To do this, let us write Ujiˆu1e1ji‡u2e2ji and determine u1andu2. To determine u1, we take the inner product of both sides with he1j: u1ˆe1hjUiˆ12p11…†1‡i 3p ‡i/C32! ˆ1 2p…1‡ 3p ‡2i†; likewise, u2ˆ1 2p…1ÿ 3p †: As a check on the calculation, let us compute the norm squared of the vector and see if it equals j1‡ij2‡j 3p ‡ij2ˆ6. We find u1jj2‡u2jj2ˆ1 2…1‡3‡2 3p ‡4‡1‡3ÿ23p †ˆ6: /C84he /C71ram/C177/C83chmidt orthogonali/C122ation process We now take up the Gram–Schmidt orthogonalization method for converting a linearly independent basis into an orthonormal one. The basic idea can be clearly illustrated in the following steps. Let j1i;j2i;...;jii;...be a linearly independent basis. To get an orthonormal basis out of these, we do the following: Step 1. Rescale the first vector by its own length, so it becomes a unit vector. This will be the first basis vector. e 1jiˆ1ji 1jijj; where 1 jijjˆ 1j1hi/C112 . Clearly e1je1 hi ˆ1j1hi 1jijjˆ1: Step 2. To construct the second member of the orthonormal basis, we subtract from the second vector j2iits projection along the first, leaving behind only the part perpendicular to the first. IIjiˆ2jiÿe1jie1j2hi : 209THE GRAM–SCHMIDT ORTHOGONALI/C90ATION PROCESS Clearly e1jIIhi ˆe1j2hi ÿe1je1hi e1j2hi ˆ0;i:e:;…IIj?je1i: Dividing jIIiby its norm (length), we now have the second basis vector and it is orthogonal to the first base vector je1iand of unit length. Step 3. To construct the third member of the orthonormal basis, consider jIIIiˆj3iÿje1ihe1jIIIiÿje2i2jIIIi which is orthogonal to both je1iandje2i. Dividing by its norm we get je3i. Continuing in this way, we will obtain an orthonormal basis je1i;je2i;...;jeni. /C84he /C67auch/C121/C177/C83ch/C119ar/C122 inequalit/C121 If/C65and/C66are non-zero vectors in /C693, then the dot product gives /C65/C66ˆABcos, where is the angle between the vectors. If we square both sides and use the fact that cos21, we obtain the inequality …/C65/C66†2A2B2or /C65/C66jj AB: This is known as the Cauchy–Schwarz inequality. There is an inequality corre- sponding to the Cauchy–Schwarz inequality in any inner-product space that obeys Axioms B.1–B.4, which can be stated as Uhj/C87i jj jUj/C87jj;Ujjˆ UhjUi/C112 etc:; …5:13† where jUiandj/C87iare two non-zero vectors in an inner-product space. This can be proved as follows. We first note that, for any scalar , the following inequality holds 0U‡ /C87 hj U‡ /C87i jj2ˆU‡ /C87 hj U‡ /C87i ˆUhjUi‡ /C87hj Ui‡Uhj /C87i‡ /C87hj /C87i ˆUjj2‡ /C42/C86hjUi‡ Uhj/C87i‡ jj2/C87jj2: Now let ˆhUj/C87i/C42=jhUj/C87ij, with real. This is possible if j/C87i6 ˆ0, but if hUj/C87iˆ0, then Cauchy–Schwarz inequality is trivial. Making this substitution in the above, we have 0Ujj2‡2Uhj/C87i jj ‡ 2/C87jj2: This is a quadratic expression in the real variable with real coecients. Therefore, the discriminant must be less than or equal to zero: 4Uhj/C87i jj2ÿ4Ujj2/C87jj20 210LINEAR VECTOR SPACES or Uhj/C87i jj Ujj/C87jj; which is the Cauchy–Schwarz inequality. From the Cauchy–Schwarz inequality follows another important inequality, known as the triangle inequality, U‡/C87 jj  Ujj ‡ /C87jj: …5:14† The proof of this is very straightforward. For any pair of vectors, we have U‡/C87 jj2ˆU‡/C87 hj U‡/C87iˆUjj2‡/C87jj2‡Uhj/C87i‡/C87hjUi Ujj2‡/C87jj2‡2Uhj/C87i jj Ujj2‡/C87jj2‡2Ujj/C87jj … Ujj2‡/C87jj2† from which it follows that U‡/C87 jj Ujj‡/C87jj: If/C86denotes the vector space of real continuous functions on the interval axb, and fand gare any real continuous functions, then the following is an inner product on /C86: fhj/C103iˆZb af…x†/C103…x†dx: The Cauchy–Schwarz inequality now gives Zb af…x†/C103…x†dx 2 Zb af2…x†dxZb a/C1032…x†dx or in Dirac notation fhj/C103ijj2fjj2/C103jj2: /C68ual /C118ectors and dual spaces We begin with a technical point regarding the inner product huj/C118i. If we set j/C118iˆ j/C119i‡/C12jzi; then huj/C118iˆ huj/C119i‡/C12hujzi is a linear function of and/C12. However, if we set juiˆ j/C119i‡/C12jzi; 211DUAL VECTORS AND DUAL SPACES then huj/C118iˆh/C118jui/C42ˆ /C42h/C118j/C119i/C42‡/C12/C42h/C118jzi/C42ˆ /C42h/C119j/C118i‡/C12/C42hzj/C118i is no longer a linear function of and /C12. To remove this asymmetry, we can introduce, besides the ket vectors ji, bra vectors hjwhich form a di/C128erent vector space. We will assume that there is a one-to-one correspondence between ket vectors ji, and bra vectors hj. Thus there are two vector spaces, the space of kets and a dual space of bras. A pair of vectors in which each is in correspondence with the other will be called a pair of dual vectors. Thus, for example, h/C118jis the dual vector of j/C118i. Note they always carry the same identification label. We now define the multiplication of ket vectors by bra vectors by requiring hujj/C118ihuj/C118i: Setting uhjˆ/C119hj /C42‡zhj/C12/C42; we have huj/C118iˆ /C42h/C119j/C118i‡/C12/C42hzj/C118i; the same result we obtained above, and we see that /C119hj /C42‡zhj/C12/C42 is the dual vector of j/C119i‡/C12jzi. From the above discussion, it is obvious that inner products are really defined only between bras and kets and hence from elements of two distinct but relatedvector spaces. There is a basis of vectors jiifor expanding kets and a similar basis hijfor expanding bras. The basis ket jtiis represented in the basis we are using by a column vector with all zeros except for a 1 in the ith row, while the basis hijis a row vector with all zeros except for a 1 in the ith column. Linear operators A useful concept in the study of linear vector spaces is that of a linear transfor- mation, from which the concept of a linear operator emerges naturally. It is instructive first to review the concept of transformation or mapping. Given vector spaces /C86and /C87and function T ~,i fT ~associates each vector in /C86with a unique vector in /C87, we say T ~maps /C86 into /C87, and write T ~:/C86!/C87.I fT ~associates the vector j/C119iin/C87with the vector j/C118iin/C86, we say that j/C119iis the image ofj/C118iunder T ~and write j/C119iˆT ~j/C118i. Further, T ~is a linear transformation if: (a)T ~…jui‡j/C118i† ˆT ~jui‡T ~j/C118ifor all vectors juiandj/C118iin/C86. (b)T ~…kj/C118i† ˆkT ~j/C118ifor all vectors j/C118iin/C86and all scalars k. We can illustrate this with a very simple example. If j/C118iˆ…x;y†is a vector in /C692, then T ~…j/C118i† ˆ … x;x‡y;xÿy†defines a function (a transformation) that maps 212LINEAR VECTOR SPACES /C692into /C693. In particular, if j/C118iˆ… 1;1†, then the image of j/C118iunder T ~isT ~…j/C118i† ˆ … 1;2;0†. It is easy to see that the transformation is linear. If juiˆ…x1;y1†andj/C118iˆ…x2;y2†, then jui‡j/C118iˆ…x1‡x2;y1‡y2†; so that T ~uji‡/C118ji …† ˆ x1‡x2;…x1‡x2†‡…y1‡y2†;…x1‡x2†ÿ…y1‡y2† …† ˆx1;x1‡y1;x1ÿy1 …† ‡ x2;x2‡y2;x2ÿy2 …† ˆT ~uji…† ‡ T ~/C118ji…† and if kis a scalar, then T ~kuji…† ˆ kx1;kx1‡ky1;kx1ÿky1 …† ˆ kx 1;x1‡y1;x1ÿy1 …† ˆ kT ~uji…†: Thus T ~is a linear transformation. IfT ~maps the vector space onto itself ( T ~:/C86!/C86), then it is called a linear operator on /C86.I n/C693a rotation of the entire space about a fixed axis is an example of an operation that maps the space onto itself. We saw in Chapter 3 that rotation can be represented by a matrix with elements ij…i;jˆ1;2;3†;i fx1;x2;x3are the components of an arbitrary vector in /C693before the transformation and x0 1;x0 2;x0 3 the components of the transformed vector, then x0 1ˆ11x1‡12x2‡13x3; x0 2ˆ21x1‡22x2‡23x3; x0 3ˆ31x1‡32x2‡33x3:9 >>= >>;…5:15† In matrix form we have ~x0ˆ~…†~x; …5:16† where is the angle of rotation, and ~x0ˆx0 1 x0 2 x0 30 B@1 CA; ~xˆx1 x2 x30 B@1 CA;and ~…†ˆ111213 212223 3132330 B@1 CA: In particular, if the rotation is carried out about x3-axis, ~…†has the following form: ~…†ˆcosÿsin0 sincos0 00 10 B@1 CA: 213LINEAR OPERATORS Eq. (5.16) determines the vector /C1200if the vector /C120is given, and ~…†is the operator (matrix representation of the rotation operator) which turns /C120into /C1200. Loosely speaking, an operator is any mathematical entity which operates on any vector in /C86and turns it into another vector in /C86. Abstractly, an operator L ~is a mapping that assigns to a vector j/C118iin a linear vector space /C86another vector jui in/C86:juiˆL ~j/C118i. The set of vectors j/C118ifor which the mapping is defined, that is, the set of vectors j/C118ifor which L ~j/C118ihas meaning, is called the domain of L ~. The set of vectors juiin the domain expressible as juiˆL ~j/C118iis called the range of the operator. An operator L ~is linear if the mapping is such that for any vectors jui;j/C119iin the domain of L ~and for arbitrary scalars ,/C12, the vector jui‡/C12j/C119iis in the domain of L ~and L ~… jui‡/C12j/C119i† ˆ L ~jui‡/C12L ~j/C119i: A linear operator is bounded if its domain is the entire space /C86and if there exists a single constant /C67such that jL ~j/C118ij<Cjj/C118ij for all j/C118iin/C86. We shall consider linear bounded operators only. Matri/C120 representation of operators Linear bounded operators may be represented by matrix. The matrix will have a finite or an infinite number of rows according to whether the dimension of /C86is finite or infinite. To show this, let j1i;j2i;...be an orthonormal basis in /C86; then every vector j’iin/C86may be written in the form j’iˆ 1j1i‡ 2j2i‡ : Since L ~jiis also in /C86, we may write L ~j’iˆ/C121j1i‡/C122j2i‡ : But L ~j’iˆ 1L ~j1i‡ 2L ~j2i‡ ; so /C121j1i‡/C122j2i‡ˆ 1L ~j1i‡ 2L ~j2i‡ : Taking the inner product of both sides with h1jwe obtain /C121ˆh1jL ~j1i 1‡h1jL ~j2i 2ˆ/C1311 1‡/C1312 2‡ ; 214LINEAR VECTOR SPACES Similarly /C122ˆh2jL ~j1i 1‡h2jL ~j2i 2ˆ/C1321 1‡/C1322 2‡ ; /C123ˆh3jL ~j1i 1‡h3jL ~j2i 2ˆ/C1331 1‡/C1332 2‡ : In general, we have /C12iˆX j/C13ij j; where /C13ijˆihjL ~jji: …5:17† Consequently, in terms of the vectors j1i;j2i;...as a basis, operator L ~is repre- sented by the matrix whose elements are /C13ij. A matrix representing L ~can be found by using any basis, not necessarily an orthonormal one. Of course, a change in the basis changes the matrix representing L ~. /C84he algebra of linear operators LetA ~andB ~be two operators defined in a linear vector space /C86of vectors ji. The equation A ~ˆB ~will be understood in the sense that A ~jiˆB ~ji for all ji2 /C86: We define the addition and multiplication of linear operators as C ~ˆA ~‡B ~and D ~ˆA ~B ~ if for any ji C ~jiˆ…A ~‡B ~†jiˆA ~ji‡B ~ji; D ~jiˆ… A ~B ~†j i ˆ A ~…B ~ji †: Note that A ~‡B ~andA ~B ~are themselves linear operators. Example 5.10 (a)…A ~B ~†… uji‡/C12/C118ji †ˆA ~‰ …B ~uji †‡/C12…B ~/C118j i†Šˆ …A ~B ~†uji‡/C12…A ~B ~†/C118ji; (b)C ~…A ~‡B ~†/C118ji ˆ C ~…A ~/C118ji ‡ B ~/C118ji † ˆ C ~A ~ji‡C ~B ~ji, which shows that C ~…A ~‡B ~†ˆC ~A ~‡C ~B ~: 215THE ALGEBRA OF LINEAR OPERATORS In general A ~B ~6ˆB ~A ~. The di/C128erence A ~B ~ÿB ~A ~is called the commutator of A ~andB ~and is denoted by the symbol ‰A ~;B ~/C93: ‰A ~;B ~ŠA ~B ~ÿB ~A ~: …5:18† An operator whose commutator vanishes is called a commuting operator. The operator equation B ~ˆ A ~ˆA ~ is equivalent to the vector equation B ~jiˆ A ~jifor any ji: And the vector equation A ~jiˆ ji is equivalent to the operator equation A ~ˆ /C69 ~ where /C69 ~is the identity (or unit) operator: /C69 ~jiˆji for any ji: It is obvious that the equation A ~ˆ is meaningless. Example 5.11 To illustrate the non-commuting nature of operators, let A ~ˆx;B ~ˆd=dx. Then A ~B ~f…x†ˆxd dxf…x†; and B ~A ~f…x†ˆd dxxf…x†ˆdx dx f‡xdf dxˆf‡xdf dxˆ…/C69 ~‡A ~B ~†f: Thus, …A ~B ~ÿB ~A ~†f…x†ˆÿ /C69 ~f…x† or x;d dx ˆxd dxÿd dxxˆÿ/C69 ~: Having defined the product of two operators, we can also define an operator raised to a certain power. For example A ~mjiˆA ~A ~ A ~ji: 216LINEAR VECTOR SPACES /C124/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C123/C122/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C125 mfactor By combining the operations of addition and multiplication, functions of opera- tors can be formed. We can also define functions of operators by their powerseries expansions. For example, eA ~formally means eA ~1‡A ~‡1 2/C33A ~2‡1 3/C33A ~3‡ : A function of a linear operator is a linear operator. Given an operator A ~that acts on vector ji, we can define the action of the same operator on vector hj. We shall use the convention of operating on hjfrom the right. Then the action of A ~on a vector hjis defined by requiring that for any jui andh/C118j, we have fuhjA ~g/C118jiuhj fA ~/C118ji gˆuhjA ~/C118ji: We may write j/C118iˆj /C118iand the corresponding bra as h /C118j. However, it is important to note that h /C118jˆA/C42h/C118j. /C69igen/C118alues and eigen/C118ectors of an operator The result of operating on a vector with an operator A ~is, in general, a di/C128erent vector. But there may be some vector j/C118iwith the property that operating with A ~on it yields the same vector j/C118imultiplied by a scalar, say : A ~j/C118iˆ j/C118i: This is called the eigenvalue equation for the operator A ~, and the vector j/C118iis called an eigenvector of A ~belonging to the eigenvalue . A linear operator has, in general, several eigenvalues and eigenvectors, which can be distinguished by asubscript A ~j/C118kiˆ kj/C118ki: The set f kgof all the eigenvalues taken together constitutes the spectrum of the operator. The eigenvalues may be discrete, continuous, or partly discrete andpartly continuous. In general, an eigenvector belongs to only one eigenvalue. If several linearly independent eigenvectors belong to the same eigenvalue, the eigen- value is said to be degenerate, and the degree of degeneracy is given by the number of linearly independent eigenvectors. /C83ome special operators Certain operators with rather special properties play very important roles inphysics. We now consider some of them below. 217EIGENVALUES AND EIGENVECTORS OF AN OPERATOR /C84he inverse of an operator The operator X ~satisfying X ~A ~ˆ/C69 ~is called the left inverse of A ~and we denote it byA ~ÿ1 L. Thus, A ~ÿ1 LA ~/C69 ~. Similarly, the right inverse of A ~is defined by the equation A ~A ~ÿ1 R/C69 ~: In general, A ~ÿ1 LorA ~ÿ1 R, or both, may not be unique and even may not exist at all. However, if both A ~ÿ1 LandA ~ÿ1 Rexist, then they are unique and equal to each other: A ~ÿ1 LˆA ~ÿ1 RA ~ÿ1; and A ~A ~ÿ1ˆA ~ÿ1A ~ˆ/C69 ~: …5:19† A ~ÿ1is called the operator inverse to A ~. Obviously, an operator is the inverse of another if the corresponding matrices are. An operator for which an inverse exists is said to be non-singular, whereas one for which no inverse exists is singular. A necessary and sucient condition for an operator A ~to be non-singular is that corresponding to each vector jui, there should be a unique vector j/C118isuch that juiˆA ~j/C118i: The inverse of a linear operator is a linear operator. The proof is simple: let ju1iˆA ~j/C1181i;ju2iˆA ~j/C1182i: Then j/C1181iˆA ~ÿ1ju1i;j/C1182iˆA ~ÿ1ju2i so that c1j/C1181iˆc1A ~ÿ1ju1i;c2j/C1182iˆc2A ~ÿ1ju2i: Thus, A ~ÿ1‰c1ju1i‡c2ju2iŠ ˆA ~ÿ1‰c1A ~j/C1181i‡c2A ~j/C1182iŠ ˆA ~ÿ1A ~‰c1j/C1181i‡c2j/C1182iŠ ˆc1j/C1181i‡c2j/C1182i or A ~ÿ1‰c1ju1i‡c2ju2iŠ ˆc1A ~ÿ1ju1i‡c2A ~ÿ1ju2i: 218LINEAR VECTOR SPACES The inverse of a product of operators is the product of the inverse in the reverse order …A ~B ~†ÿ1ˆB ~ÿ1A ~ÿ1: …5:20† The proof is straightforward: we have A ~B ~…A ~B ~†ÿ1ˆ/C69 ~: Multiplying successively from the left by A ~ÿ1andB ~ÿ1, we obtain …A ~B ~†ÿ1ˆB ~ÿ1A ~ÿ1; which is identical to Eq. (5.20). /C84he ad/C106oint operators Assuming that /C86is an inner-product space, then the operator X ~satisfying the relation uhjX ~/C118jiˆ/C118hjA ~uji/C42 for any jui;j/C118i2/C86 is called the adjoint operator of A ~and is denoted by A ~‡. Thus uhjA ~‡/C118ji/C118hjA ~uji/C42 for any jui;j/C118i2/C86: …5:21† We first note that hjA ~‡is a dual vector of A ~ji. Next, it is obvious that …A ~‡†‡ˆA ~: …5:22†: To see this, let A ~‡ˆB ~, then ( A ~‡†‡becomes B ~‡, and from Eq. (5.21) we find /C118hjB ~‡ujiˆuhjB ~/C118ji/C42;for any jui;j/C118i2/C86: But uhjB ~/C118ji/C42ˆuhjA ~‡/C118ji/C42ˆ/C118hjA ~uji: Thus /C118hjB ~‡ujiˆuhjB ~/C118ji/C42ˆ/C118hjA ~uji from which we find …A ~‡†‡ˆA ~: 219SOME SPECIAL OPERATORS It is also easy to show that …A ~B ~†‡ˆB ~‡A ~‡: …5:23† For any jui;j/C118i,h/C118jB ~‡andB ~j/C118iis a pair of dual vectors; hujA ~‡andA ~juiis also a pair of dual vectors. Thus we have /C118hjB ~‡A ~‡uji ˆ f /C118hjB ~‡gfA ~‡uji g ˆ ‰ f uhjA ~gfB ~/C118j igŠ/C42 ˆuhjA ~B ~/C118ji/C42ˆ/C118hj …A ~B ~†‡uji and therefore …A ~B ~†‡ˆB ~‡A ~‡: /C72ermitian operators An operator H ~that is equal to its adjoint, that is, that obeys the relation H ~ˆH ~‡…5:24† is called Hermitian or self-adjoint. And H ~is anti-Hermitian if H ~ˆÿH ~‡: Hermitian operators have the following important properties: (1) The eigenvalues are real: Let H ~be the Hermitian operator and let j/C118ibe an eigenvector belonging to the eigenvalue : H ~j/C118iˆ j/C118i: By definition, we have /C118hjA ~/C118jiˆ/C118hjA ~/C118ji/C42; that is, … /C42ÿ †/C118/C118jih ˆ0: Since h/C118j/C118i6 ˆ0, we have /C42ˆ : (2) Eigenvectors belonging to di/C128erent eigenvalues are orthogonal: Let juiand j/C118ibe eigenvectors of H ~belonging to the eigenvalues and/C12respectively: H ~juiˆ jui;H ~j/C118iˆ/C12j/C118i: 220LINEAR VECTOR SPACES Then uhjH ~/C118jiˆ/C118hjH ~uji/C42: That is, … ÿ/C12†h/C118juiˆ0…since /C42ˆ †: But 6ˆ/C12, so that h/C118juiˆ0: (3) The set of all eigenvectors of a Hermitian operator forms a complete set: The eigenvectors are orthogonal, and since we can normalize them, this means that the eigenvectors form an orthonormal set and serve as a basis for the vector space. /C85nitary operators A linear operator U ~is unitary if it preserves the Hermitian character of an operator under a similarity transformation: …U ~A ~U ~ÿ1†‡ˆU ~A ~U ~ÿ1; where A ~‡ˆA ~: But, according to Eq. (5.23) …U ~A ~U ~ÿ1†‡ˆ…U ~ÿ1†‡A ~U ~‡; thus, we have …U ~ÿ1†‡A ~U ~‡ˆU ~A ~U ~ÿ1: Multiplying from the left by U ~‡and from the right by U ~, we obtain U ~‡…U ~ÿ1†‡A ~U ~‡U ~ˆU ~‡U ~A ~; this reduces to A ~…U ~‡U ~†ˆ…U ~‡U ~†A ~; since U ~‡…U ~ÿ1†‡ˆ…U ~ÿ1U ~†‡ˆ/C69 ~: Thus U ~‡U ~ˆ/C69 ~ 221SOME SPECIAL OPERATORS or U ~‡ˆU ~ÿ1: …5:25† We often use Eq. (5.25) for the definition of the unitary operator. Unitary operators have the remarkable property that transformation by a uni- tary operator preserves the inner product of the vectors. This is easy to see: under the operation U ~, a vector j/C118iis transformed into the vector j/C1180iˆU ~j/C118i. Thus, if two vectors j/C118iandjuiare transformed by the same unitary operator U ~, then u0/C10/C12/C12/C1180/C11 ˆhU ~ujU ~/C118iˆuhjU ~‡U ~/C118iˆu/C118jih ; that is, the inner product is preserved. In particular, it leaves the norm of a vector unchanged. Thus, a unitary transformation in a linear vector space is analogous to a rotation in the physical space (which also preserves the lengths of vectors and the inner products). Corresponding to every unitary operator U ~, we can define a Hermitian opera- torH ~and vice versa by U ~ˆei/C34H ~; …5:26† where /C34is a parameter. Obviously U ~‡ˆe…i/C34H ~†‡ =eÿi/C34H ~ˆU ~ÿ1: A unitary operator possesses the following properties: (1) The eigenvalues are unimodular; that is, if U ~j/C118iˆ j/C118i, then j jˆ1. (2) Eigenvectors belonging to di/C128erent eigenvalues are orthogonal.(3) The product of unitary operators is unitary. /C84he pro/C106ection operators A symbol of the type of juih/C118jis quite useful: it has all the properties of a linear operator, multiplied from the right by a ket ji, it gives juiwhose magnitude is h/C118ji; and multiplied from the left by a bra hjit gives h/C118jwhose magnitude is hjui. The linearity of juih/C118jresults from the linear properties of the inner product. We also have fuji/C118hj g ‡ˆ/C118jiuhj: The operator P ~jˆjjijhjis a very particular example of projection operator. To see its e/C128ect on an arbitrary vector jui, let us expand jui: 222LINEAR VECTOR SPACES uji ˆXn jˆ1ujjji;ujˆjhjui: …5:27† We may write the above as ujiˆXn jˆ1jjijhj/C32! uji; which is true for all jui. Thus the object in the brackets must be identified with the identity operator: I ~ˆXn jˆ1jjijhjˆXn jˆ1P ~j: …5:28† Now we will see that the e/C128ect of this particular projection operator on juiis to produce a new vector whose direction is along the basis vector jjiand whose magnitude is hjjui: P ~jujiˆjjijhjuiˆjjiuj: We see that whatever juiis,P ~jjuiis a multiple of jjiwith a coecient ujwhich is the component of juialong jji. Eq. (5.28) says that the sum of the projections of a vector along all the ndirections equals the vector itself. When P ~jˆjjijhjacts on jji, it reproduces that vector. On the other hand, since the other basis vectors are orthogonal to jji, a projection operation on any one of them gives zero (null vector). The basis vectors are therefore eigenvectors of P ~k with the property P ~kjjiˆkjjji; …j;kˆ1;...;n†: In this orthonormal basis the projection operators have the matrix form P ~1ˆ100  000  000  ............0 BBBB@1 CCCCA;P ~2ˆ000  010  000  ............0 BBBB@1 CCCCA;P ~/C78ˆ000  000  000  ............ 10 BBBBB@1 CCCCCA: Projection operators can also act on bras in the same way: uhjP ~jˆuhjjijhjˆuj/C42jhj: 223SOME SPECIAL OPERATORS /C67hange of basis The choice of base vectors (basis) is largely arbitrary and di/C128erent representations are physically equally acceptable. How do we change from one orthonormal set of base vectors j’1i;j’2i;...;j/C34nito another such set j/C241i;j/C242i;...;j/C24ni/C63 In other words, how do we generate the orthonomal set j/C241i;j/C242i;...;j/C24nifrom the old setj’1i;j’2i;...;j’n/C63 This task can be accomplished by a unitary transformation: j/C24iiˆU ~j’ii…iˆ1;2;...;n†: …5:29† Then given a vector XjiˆPn iˆ1ai’iji, it will be transformed into jX0i: jX0iˆU ~jXiˆU ~Xn iˆ1ai’ijiˆXn iˆ1U ~ai’ijiˆXn iˆ1ai/C24iji: We can see that the operator U ~possesses an inverse U ~ÿ1which is defined by the equation j’iiˆU ~ÿ1j/C24ii…iˆ1;2;...;n†: The operator U ~is unitary; for, if XjiˆPniˆ1ai’ijiand YjiˆPniˆ1bi’iji, then XjYhi ˆXn i;jˆ1ai/C42bj’ij’j/C10/C11 ˆXn iˆ1ai/C42bi; hUXjUYiˆXn i;jˆ1ai/C42bj/C24ij/C24j/C10/C11 ˆXn iˆ1ai/C42bi: Hence Uÿ1ˆU‡: The inner product of two vectors is independent of the choice of basis which spans the vector space, since unitary transformations leave all inner products invariant. In quantum mechanics inner products give physically observable quan- tities, such as expectation values, probabilities, etc. It is also clear that the matrix representation of an operator is di/C128erent in a di/C128erent basis. To find the e/C128ect of a change of basis on the matrix representation of an operator, let us consider the transformation of the vector jXiintojYiby the operator A ~: jYiˆA ~j…Xi: …5:30† Referred to the basis j’1i;j’2i;...;j’i;jXiand jYiare given by jXiˆPn iˆ1ai’iijandjYiˆPniˆ1bij’ii, and the equation jYiˆA ~jXibecomes Xn iˆ1bi’ijiˆA ~Xn jˆ1aj’j/C12/C12/C11 : 224LINEAR VECTOR SPACES Multiplying both sides from the left by the bra vector h’ijwe find biˆXn jˆ1aj’ihjA ~’j/C12/C12/C11 ˆXn jˆ1ajAij: …5:31† Referred to the basis j/C241i;j/C242i;...;j/C24nithe same vectors jXiand jYiare jXiˆPn iˆ1a0 i/C24iij, and jYiˆPniˆ1b0 ij/C24ii, and Eqs. (5.31) are replaced by b0 iˆXn jˆ1a0 j/C24ihjA ~/C24j/C12/C12/C11 ˆXn jˆ1a0 jA0 ij; where A0 ijˆh/C24ijA ~/C24ji/C12/C12, which is related to A ijby the following relation: A0 ijˆ/C24ihjA ~/C24j/C12/C12/C11 ˆU’ihj A ~U’j/C12/C12/C11 ˆ’ ihjU/C42A ~U’j/C12/C12/C11 ˆ…U/C42A ~U†ij or using the rule for matrix multiplication A0 ijˆ/C24ihjA ~/C24j/C12/C12/C11 ˆ…U/C42A ~U†ijˆXn rˆ1Xn sˆ1Uir/C42ArsUsj: …5:32† From Eqs. (5.32) we can find the matrix representation of an operator with respect to a new basis. If the operator A ~transforms vector jXiinto vector jYiwhich is vector jXiitself multiplied by a scalar /C58jYiˆjXi, then Eq. (5.30) becomes an eigenvalue equation: A ~jXiˆjXi: /C67ommuting operators In general, operators do not commute. But commuting operators do exist andthey are of importance in quantum mechanics. As Hermitian operators play a dominant role in quantum mechanics, and the eigenvalues and the eigenvectors of a Hermitian operator are real and form a complete set, respectively, we shall concentrate on Hermitian operators. It is straightforward to prove that Two commuting Hermitian operators possess a complete ortho-normal set of common eigenvectors, and vice versa. IfA ~andA ~j/C118iˆ j/C118iare two commuting Hermitian operators, and if A ~j/C118iˆ j/C118i; …5:33† then we have to show that B ~j/C118iˆ/C12j/C118i: …5:34† 225COMMUTING OPERATORS Multiplying Eq. (5.33) from the left by B ~, we obtain B ~…A ~j/C118i† ˆ …B ~j/C118i†; which using the fact A ~B ~ˆB ~A ~, can be rewritten as A ~…B ~j/C118i† ˆ …B ~j/C118i†: Thus, B ~j/C118iis an eigenvector of A ~belonging to eigenvalue .I f is non-degen- erate, then B ~j/C118ishould be linearly dependent on j/C118i, so that a…B ~j/C118i† ‡bj/C118iˆ0;with a6ˆ0 and b6ˆ0: It follows that B ~j/C118iˆÿ … b=a†j/C118iˆ/C12j/C118i: IfAis degenerate, then the matter becomes a little complicated. We now state the results without proof. There are three possibilities: (1) The degenerate eigenvectors (that is, the linearly independent eigenvectors belonging to a degenerate eigenvalue) of A ~are degenerate eigenvectors of B ~also. (2) The degenerate eigenvectors of A ~belong to di/C128erent eigenvalues of B ~. In this case, we say that the degeneracy is removed by the Hermitian operator B ~. (3) Every degenerate eigenvector of A ~is not an eigenvector of B ~. But there are linear combinations of the degenerate eigenvectors, as many in number as the degrees of degeneracy, which are degenerate eigenvectors of A ~but are non-degenerate eigenvectors of B ~. Of course, the degeneracy is removed byB ~. Function spaces We have seen that functions can be elements of a vector space. We now return to this theme for a more detailed analysis. Consider the set of all functions that arecontinuous on some interval. Two such functions can be added together to con- struct a third function /C104…x†: /C104…x†ˆf…x†‡/C103…x†;axb; where the plus symbol has the usual operational meaning of ‘add the value of fat the point xto the value of gat the same point.’ A function f…x†can also be multiplied by a number kto give the function /C112…x†: /C112…x†ˆkf…x†;axb: 226LINEAR VECTOR SPACES The centred dot, the multiplication symbol, is again understood in the conven- tional meaning of ‘multiply by kthe value of f…x†at the point x.’ It is evident that the following conditions are satisfied: (a) By adding two continuous functions, we obtain a continuous function. (b) The multiplication by a scalar of a continuous function yields again a con- tinuous function. (c) The function that is identically zero for axbis continuous, and its addition to any other function does not alter this function. (d) For any function f…x†there exists a function …ÿ1†f…x†, which satisfies f…x† ‡ ‰…ÿ 1†f…x†Š ˆ0: Comparing these statements with the axioms for linear vector spaces (Axioms A.1–A.8), we see clearly that the set of all continuous functions defined on someinterval forms a linear vector space; this is called a function space. We shall consider the entire set of values of a function f…x†as representing a vector jfi of this abstract vector space /C70(/C70stands for function space). In other words, we shall treat the number f…x†at the point xas the component with ‘index x’o fa n abstract vector jfi. This is quite similar to what we did in the case of finite- dimensional spaces when we associated a component a iof a vector with each value of the index i. The only di/C128erence is that this index assumed a discrete set of values 1, 2, etc., up to /C78(for/C78-dimensional space), whereas the argument xof a function f…x†is a continuous variable. In other words, the function f…x†has an infinite number of components, namely the values it takes in the continuum of points labeled by the real variable x. However, two questions may be raised. The first question concerns the orthonormal basis. The components of a vector are defined with respect to some basis and we do not know which basis has been(or could be) chosen in the function space. Unfortunately, we have to postpone the answer to this question. Let us merely note that, once a basis has been chosen, we work only with the components of a vector. Therefore, provided we do not change to other basis vectors, we need not be concerned about the particular basis that has been chosen. The second question is how to define an inner product in an infinite-dimen- sional vector space. Suppose the function f…x†describes the displacement of a string clamped at xˆ0 and xˆL. We divide the interval of length Linto /C78equal parts and measure the displacements f…x i†fiat/C78point xi;iˆ1;2;...;/C78.A t fixed /C78, the functions are elements of a finite /C78-dimensional vector space. An inner product is defined by the expression fhj/C103iˆX/C78 iˆ1fi/C103i: 227FUNCTION SPACES For a vibrating string, the space is real and there is no need to conjugate anything. To improve the description, we can increase the number /C78. However, as /C78!1 by increasing the number of points without limit, the inner product diverges as we subdivide further and further. The way out of this is to modify the definition by a positive prefactor ˆL=/C78which does not violate any of the axioms for the inner product. But now fhj/C103iˆlim !0X/C78 iˆ1fi/C103i!ZL 0f…x†/C103…x†dx; by the usual definition of an integral. Thus the inner product of two functions is the integral of their product. Two functions are orthogonal if this inner product vanishes, and a function is normalized if the integral of its square equals unity. Thus we can speak of an orthonormal set of functions in a function space just as in finite dimensions. The following is an example of such a set of functions defined in the interval 0 xLand vanishing at the end points: emji!m…x†ˆ 2 L/C114 sinmx L; mˆ1;2;...;1; emhjeniˆ2 LZL 0sinmx Lsinnx Ldxˆmn: For the details, see ‘Vibrating strings’ of Chapter 4. In quantum mechanics we often deal with complex functions and our definition of the inner product must then modified. We define the inner product of f…x†and /C103…x†as fhj/C103iˆZL 0f/C42…x†/C103…x†dx; where f/C42 is the complex conjugate of f. An orthonormal set for this case is m…x†ˆ1  2peimx; mˆ0;1;2;...; which spans the space of all functions of period 2 with finite norm. A linear vector space with a complex-type inner product is called a /C72ilbert space . Where and how did we get the orthonormal functions/C63 In general, by solving the eigenvalue equation of some Hermitian operator. We give a simple example here. Consider the derivative operator Dˆd…†=dx/C58 Df…x†ˆdf…x†=dx;Djfiˆdjfi=dx: However, /C68is not Hermitian, because it does not meet the condition: ZL 0f/C42…x†d/C103…x† dxdxˆZL 0/C103/C42…x†df…x† dxdx/C42: 228LINEAR VECTOR SPACES Here is why: ZL 0/C103/C42…x†df…x† dxdx/C42ˆZL 0/C103…x†df/C42…x† dxdx ˆ/C103f/C42L 0ÿZL 0f/C42…x†d/C103…x† dxdx:/C12/C12/C12/C12/C12 It is easy to see that hermiticity of /C68is lost on two counts. First we have the term coming from the end points. Second the integral has the wrong sign. We can fix both of these by doing the following: (a) Use operator ÿiD. The extra iwill change sign under conjugation and kill the minus sign in front of the integral. (b) Restrict the functions to those that are periodic: f…0†ˆf…L†. Thus, ÿiDis a Hermitian operator on period functions. Now we have ÿidf…x† dxˆf…x†; where is the eigenvalue. Simple integration gives f…x†ˆAeix: Now the periodicity requirement gives eiLˆei0ˆ1 from which it follows that ˆ2m=L; mˆ0;1;2; and the normalization condition gives Aˆ1 Lp: Hence the set of orthonormal eigenvectors is given by fm…x†ˆ1Lpe2imx=L: In quantum mechanics the eigenvalue equation is the Schro /C200dinger equation and the Hermitian operator is the Hamiltonian operator. Quantum mechanically, a system with ndegrees of freedom which is classically specified by ngeneralized coordinates /C1131;...;/C1132;/C113nis specified at a fixed instant of time by a wave function /C32…/C1131;/C1132;...;/C113n†whose norm is unity, that is, /C32hj/C32iˆZ /C32…/C1131;/C1132;...;/C113n† jj2d/C1131;d/C1132;...;d/C113nˆ1; 229FUNCTION SPACES the integration being over the accessible values of the coordinates /C1131;/C1132;...;/C113n. The set of all such wave functions with unit norm spans a Hilbert space /C72. Every possible state of the system is represented by a function in this Hilbert space, and conversely, every vector in this Hilbert space represents a possible state of the system. In addition to depending on the coordinates /C1131;/C1132;...;/C113n, the wave func- tion depends also on the time t, but the dependence on the /C113s and on tare essentially di/C128erent. The Hilbert space H is formed with respect to the spatial coordinates /C1131;/C1132;...;/C113nonly, for example, the inner product is formed with respect to the /C113s only, and one wave function /C32…/C1131;/C1132;...;/C113n) states its complete spatial dependence. On the other hand the states of the system at di/C128erent instants of time t1;t2;... are given by the di/C128erent wave functions /C321…/C1131;/C1132;...;/C113n†;/C322…/C1131;/C1132;...;/C113n†...of the Hilbert space. Problems 5.1 Prove the three main properties of the dot product given by Eq. (5.7). 5.2 Show that the points on a line /C86passing through the origin in /C693form a linear vector space under the addition and scalar multiplication operationsfor vectors in /C69 3. Hint: The points of /C86satisfy parametric equations of the form x1ˆat;x2ˆbt;x3ˆct; ÿ1 <t<1: 5.3 Do all Hermitian 2 2 matrices form a vector space under addition/C63 Is there any requirement on the scalars that multiply them/C63 5.4 Let /C86be the set of all points ( x1;x2)i n/C692that lie in the first quadrant; that is, such that x10 and x20. Show that the set /C86fails to be a vector space under the operations of addition and scalar multiplication.Hint: Consider u/C61(1, 1) which lies in /C86. Now form the scalar multiplication …ÿ1†uˆ… ÿ 1;ÿ1†; where is this point located/C63 5.5 Show that the set /C87of all 2 2 matrices having zeros on the principal diagonal is a subspace of the vector space M 22of all 2 2 matrices. 5.6 Show that j/C87iˆ… 4;ÿ1;8†is not a linear combination of jUiˆ… 1;2;ÿ1† andj/C86iˆ… 6;4;2†. 5.7 Show that the following three vectors in /C693cannot serve as base vectors of /C693: 1jiˆ…1;1;2†;2jiˆ…1;0;1†;and 3 jiˆ…2;1;3†: 5.8 Determine which of the following lie in the space spanned by jfiˆcos2x andj/C103iˆsin2x:(a) cos 2 x;…b†3‡x2;…c†1;…d†sinx. 5.9 Determine whether the three vectors 1jiˆ…1;ÿ2;3†;2jiˆ…5;6;ÿ1†;3jiˆ…3;2;1† are linearly dependent or independent. 230LINEAR VECTOR SPACES 5.10 Given the following three vectors from the vector space of real 2 2 matrices: 1jiˆ01 00 ;2jiˆ1101 ;3jiˆÿ2ÿ1 0ÿ2 ; determine whether they are linearly dependent or independent. 5.11 If Sˆ1ji;2ji;...;nji fg is a basis for a vector space /C86, show that every set with more than nvectors is linearly dependent. 5.12 Show that any two bases for a finite-dimensional vector space have the same number of vectors. 5.13 Consider the vector space /C69 3with the Euclidean inner product. Apply the Gram–Schmidt process to transform the basis j1iˆ… 1;1;1†;j2iˆ… 0;1;1†;j3iˆ… 0;0;1† into an orthonormal basis. 5.14 Consider the two linearly independent vectors of Example 5.10: jUiˆ… 3ÿ4i†j1i‡…5ÿ6i†j2i; j/C87iˆ… 1ÿi†j1i‡…2ÿ3i†j2i; where j1iandj2iare an orthonormal basis. Apply the Gram–Schmidt pro- cess to transform the two vectors into an orthonormal basis. 5.15 Show that the eigenvalue of the square of an operator is the square of the eigenvalue of the operator. 5.16 Show that if, for a given A ~, both operators A ~ÿ1 LandA ~ÿ1 Rexist, then A ~ÿ1 LˆA ~ÿ1 RA ~ÿ1: 5.17 Show that if a unitary operator U ~can be written in the form U ~ˆ1‡ie/C70 ~, where eis a real infinitesimally small number, then the operator /C70 ~is Hermitian. 5.18 Show that the di/C128erential operator /C112 ~ˆp id dx is linear and Hermitian in the space of all di/C128erentiable wave functions /C30…x† that, say, vanish at both ends of an interval ( a,b). 5.19 The translation operator T…a†is defined to be such that T…a†/C30…x†ˆ /C30…x‡a†. Show that: (a)T…a†may be expressed in terms of the operator /C112 ~ˆp id dx; 231PROBLEMS (b)T…a†is unitary. 5.21 Verify that: …a†2 LZL 0sinmx Lsinnx Ldxˆmn: …b†1  2pZ2 0ei…mÿn†dxˆmn: 232LINEAR VECTOR SPACES 6 /C70unctions of a complex variable The theory of functions of a complex variable is a basic part of mathematical analysis. It provides some of the very useful mathematical tools for physicists and engineers. In this chapter a brief introduction to complex variables is presented which is intended to acquaint the reader with at least the rudiments of this important subject. /C67omple/C120 numbers The number system as we know it today is a result of gradual development. The natural numbers (positive integers 1, 2, ...) were first used in counting. Negative integers and zero (that is, 0, ÿ1;ÿ2;...) then arose to permit solutions of equa- tions such as x‡3ˆ2. In order to solve equations such as bxˆafor all integers aandbwhere b6ˆ0, rational numbers (or fractions) were introduced. Irrational numbers are numbers which cannot be expressed as a/C47b,w i t h aandbintegers and b6ˆ0, such as 2p ˆ1:41423 ;ˆ3:14159 Rational and irrational numbers are all real numbers. However, the real num- ber system is still incomplete. For example, there is no real number xwhich satisfies the algebraic equation x2‡1ˆ0/C58xˆ ÿ1p . The problem is that we do not know what to make of  ÿ1p because there is no real number whose square isÿ1. Euler introduced the symbol iˆ  ÿ1p in 1777 years later Gauss used the notation a‡ibto denote a complex number, where aandbare real numbers. Today, iˆ ÿ1p is called the unit imaginary number. In terms of i, the answer to equation x 2‡1ˆ0i sxˆi. It is postulated that i will behave like a real number in all manipulations involving addition and multi- plication. We now introduce a general complex number, in Cartesian form zˆx‡iy …6:1† 233 and refer to xandyas its real and imaginary parts and denote them by the symbols Re zand Im z, respectively. Thus if zˆÿ3‡2i, then Re zˆÿ3 and Imzˆ‡2. A number with just y6ˆ0 is called a pure imaginary number. The complex conjugate, or briefly conjugate, of the complex number zˆx‡iy is z/C42ˆxÿiy …6:2† and is called ‘ z-star’. Sometimes we write it zand call it ‘ z-bar’. Complex con- jugation can be viewed as the process of replacing ibyÿiwithin the complex number. Basic operations /C119ith complex numbers Two complex numbers z1ˆx1‡iy1andz2ˆx2‡iy2are equal if and only if x1ˆx2andy1ˆy2. In performing operations with complex numbers we can proceed as in the algebra of real numbers, replacing i2byÿ1 when it occurs. Given two complex numbers z1andz2where z1ˆa‡ib;z2ˆc‡id, the basic rules obeyed by com- plex numbers are the following: (1) Addition: z1‡z2ˆ…a‡ib†‡…c‡id†ˆ… a‡c†‡i…b‡d†: (2) Subtraction: z1ÿz2ˆ…a‡ib†ÿ…c‡id†ˆ… aÿc†‡i…bÿd†: (3) Multiplication: z1z2ˆ…a‡ib†…c‡id†ˆ… acÿbd†‡i…adÿbc†: (4) Division: z1 z2ˆa‡ib c‡idˆ…a‡ib†…cÿid† …c‡id†…cÿid†ˆac‡bd c2‡d2‡ibcÿad c2‡d2: Polar form of complex numbers All real numbers can be visualized as points on a straight line (the x-axis). A complex number, containing two real numbers, can be represented by a point in a two-dimensional xyplane, known as the zplane or the complex plane (also known as the Gauss plane or Argand diagram). The complex variablezˆx‡iyand its complex conjugation z/C42 are labeled in Fig. 6.1. 234FUNCTIONS OF A COMPLE/C88 VARIABLE The complex variable can also be represented by the plane polar coordinates (r;): zˆr…cos‡isin†: With the help of Euler’s formula eiˆcos‡isin; we can rewrite the last equation in polar form: zˆr…cos‡isin†ˆrei;rˆ x2‡y2q ˆ zz/C42p : …6:3† ris called the modulus or absolute value of z, denoted by jzjor mod z; and is called the phase or argument of zand it is denoted by arg z. For any complex number z6ˆ0 there corresponds only one value of in 02. The absolute value of zhas the following properties. If z1;z2;...;zmare complex numbers, then we have: (1)jz1z2zmjˆjz1jjz2jjzmj: (2)z1 z2/C12/C12/C12/C12/C12/C12/C12/C12ˆjz1j jz2j;z26ˆ0. (3)jz1‡z2‡‡ zmjjz1j‡jz2j‡‡j zmj. (4)jz1z2jjz1jÿjz2j. Complex numbers zˆreiwith rˆ1 have jzjˆ1 and are called unimodular. 235COMPLE/C88 NUMBERS Figure 6.1. The complex plane. We may imagine them as lying on a circle of unit radius in the complex plane. Special points on this circle are ˆ0…1† ˆ=2…i† ˆ…ÿ1† ˆÿ=2…ÿi†: The reader should know these points at all times. Sometimes it is easier to use the polar form in manipulations. For example, to multiply two complex numbers, we multiply their moduli and add their phases; todivide, we divide by the modulus and subtract the phase of the denominator: zz 1ˆ…rei†…r1ei1†ˆrr1ei…‡1†;z z1ˆrei r1ei1ˆr r1ei…ÿ1†: On the other hand to add two complex numbers we have to go back to theCartesian forms, add the components and revert to the polar form. If we view a complex number zas a vector, then the multiplication of zbye i (where is real) can be interpreted as a rotation of zcounterclockwise through angle ; and we can consider ei as an operator which acts on zto produce this rotation. Similarly, the multiplication of two complex numbers represents a rota-tion and a change of length: z 1ˆr1ei1;z2ˆr2ei2,z1z2ˆr1r2ei…1‡2†; the new complex number has length r1r2and phase 1‡2. Example 6.1Find …1‡i† 8. Solution: We first write zin polar form: zˆ1‡iˆr…cos‡isin†, from which we find rˆ 2p ;ˆ=4. Then zˆ 2p cos=4‡isin=4 …† ˆ2p ei=4: Thus …1‡i†8ˆ…2p ei=4†8ˆ16e2iˆ16: Example 6.2 Show that 1‡ 3p i 1ÿ 3p i/C32!10 ˆÿ1 2‡i 3p 2: 236FUNCTIONS OF A COMPLE/C88 VARIABLE 1‡i 3p 1ÿi3p/C32! 10 ˆ2ei=3 2eÿi=3/C32!10 ˆe2i=310 ˆe20i=3 ˆe6ie2i=3ˆ1c o s…2=3†‡isin…2=3† ‰Š ˆ ÿ1 2‡i 3p 2: /C68e /C77oivre’s theorem and roots of complex numbers Ifz1ˆr1ei1andz2ˆr2ei2, then z1z2ˆr1r2ei…1‡2†ˆr1r2‰cos…1‡2†‡isin…1‡2†Š: A generalization of this leads to z1z2znˆr1r2rnei…1‡2‡‡ n† ˆr1r2rn‰cos…1‡2‡‡ n†‡isin…1‡2‡‡ n†Š; ifz1ˆz2ˆˆ znˆzthis becomes znˆ…rei†nˆrn‰cos…n†‡isin…n†Š; from which it follows that …cos‡isin†nˆcos…n†‡isin…n†; …6:4† a result known as De Moivre’s theorem. Thus we now have a general rule for calculating the nth power of a complex number z. We first write zin polar form zˆr…cos‡isin†, then znˆrn…cos‡isin†nˆrn‰cosn‡isinnŠ: …6:5† The general rule for calculating the nth root of a complex number can now be derived without diculty. A number wis called an nth root of a complex number zif/C119nˆz, and we write /C119ˆz1=n.I fzˆr…cos‡isin†, then the complex num- ber /C1190ˆrnpcos n‡isin n is definitely the nth root of zbecause /C119n 0ˆz. But the numbers /C119kˆrnpcos‡2k n‡isin‡2k n ; kˆ1;2;...;…nÿ1†; are also nth roots of zbecause /C119n kˆz. Thus the general rule for calculating the nth root of a complex number is /C119ˆrnpcos‡2k n‡isin‡2k n ; kˆ0;1;2;...;…nÿ1†: …6:6† 237COMPLE/C88 NUMBERS It is customary to call the number corresponding to kˆ0 (that is, /C1190) the princi- pal root of z. Thenth roots of a complex number zare always located at the vertices of a regular polygon of nsides inscribed in a circle of radiusrnpabout the origin. Example 6.3 Find the cube roots of 8. Solution: In this case zˆ8‡i0ˆr…cos‡isin†;rˆ2 and the principal argu- ment ˆ0. Formula (6.6) then yields  83p ˆ2 cos2k 3‡isin2k 3 ; kˆ0;1;2: These roots are plotted in Fig. 6.2: 2 …kˆ0;ˆ08†; ÿ1‡i 3p …kˆ1;ˆ120 8†; ÿ1ÿi 3p …kˆ2;ˆ240 8†: Functions of a comple/C120 /C118ariable Complex numbers zˆx‡iybecome variables if xory(or both) vary. Then functions of a complex variable may be formed. If to each value which a complex variable zcan assume there corresponds one or more values of a complex variable w, we say that wis a function of zand write /C119ˆf…z†or/C119ˆ/C103…z†, etc. The variable zis sometimes called an independent variable, and then wis a dependent 238FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.2. The cube roots of 8. variable. If only one value of wcorresponds to each value of z, we say that wis a single-valued function of zor that f…z†is single-valued; and if more than one value ofwcorresponds to each value of z,wis then a multiple-valued function of z. For example, /C119ˆz2is a single-valued function of z, but /C119ˆ zpis a double-valued function of z. In this chapter, whenever we speak of a function we shall mean a single-valued function, unless otherwise stated. Mapping Note that wis also a complex variable and so can be written in the form /C119ˆu‡i/C118ˆf…x‡iy†; …6:7† where uand/C118are real. By equating real and imaginary parts this is seen to be equivalent to uˆu…x;y†;/C118ˆ/C118…x;y†: …6:8† If/C119ˆf…z†is a single-valued function of z, then to each point of the complex z plane, there corresponds a point in the complex wplane. If f…z†is multiple-valued, a point in the zplane is mapped in general into more than one point. The following two examples show the idea of mapping clearly. Example 6.4 Map /C119ˆz2ˆr2e2i: Solution: This is single-valued function. The mapping is unique, but not one-to-one. It is a two-to-one mapping, since zandÿzgive the same square. For example as shown in Fig. 6.3, zˆÿ2‡iandzˆ2ÿiare mapped to the same point wˆ3ÿ4i;a n d zˆ1ÿ3iandÿ1‡3iare mapped into the same point wˆÿ8ÿ6i. The line joining the points P…ÿ2;1†and/C81…1;ÿ3†in the z-plane is mapped by /C119ˆz2into a curve joining the image points P0…3;ÿ4†and/C810…ÿ8;ÿ6†. It is not 239MAPPING Figure 6.3. The mapping function /C119ˆz2. very dicult to determine the equation of this curve. We first need the equation of the line joining Pand/C81in the zplane. The parametric equations of the line joining Pand/C81are given by xÿ… ÿ 2† 1ÿ… ÿ 2†ˆyÿ1 ÿ3ÿ1ˆtor xˆ3tÿ2;yˆ1ÿ4t: The equation of the line P/C81is then given by zˆ3tÿ2‡i…1ÿ4t†. The curve in thewplane into which the line P/C81is mapped has the equation /C119ˆz2ˆ‰3tÿ2‡i…1ÿ4t†Š2ˆ3ÿ4tÿ7t2‡i…ÿ4‡22tÿ24t2†; from which we obtain uˆ3ÿ4tÿ7t2; /C118ˆÿ4‡22tÿ24t2: By assigning various values to the parameter t, this curve may be graphed. Sometimes it is convenient to superimpose the zandwplanes. Then the images of various points are located on the same plane and the function /C119ˆf…z†may be said to transform the complex plane to itself (or a part of itself). Example 6.5 Map /C119ˆf…z†ˆ zp;zˆrei: Solution: There are two square roots: f1…rei†ˆrpei=2;f2ˆÿf1ˆrpei…‡2†=2. The function is double-valued, and the mapping is one-to-two. This is shown in Fig. 6.4, where for simplicity we have used the same complex plane for both zand wˆf…z†. /C66ranch lines and /C82iemann surfaces We now take a close look at the function /C119ˆ zpof Example 6.5. Suppose we allow z to make a complete counterclockwise motion around the origin starting from point 240FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.4. The mapping function /C119ˆ zp: A, as shown in Fig. 6.5. At A,ˆ1and/C119ˆrpei=2. After a complete circuit back toA;ˆ1‡2and/C119ˆrpei…‡2†=2ˆÿrpei=2. However, by making a second complete circuit back to A,ˆ1‡4, and so /C119ˆrpei…‡4†=2ˆrpei=2; that is, we obtain the same value of wwith which we started. We can describe the above by stating that if 0 <2we are on one branch of the multiple-valued function zp, while if 2 <4we are on the other branch of the function. It is clear that each branch of the function is single-valued. In order to keep the function single-valued, we set up an artificial barrier such as OB (the wavy line) which we agree not to cross. This artificial barrier is called abranch line or branch cut, point Ois called a branch point. Any other line from Ocan be used for a branch line. Riemann (George Friedrich Bernhard Riemann, 1826–1866) suggested another way to achieve the purpose of the branch line described above. Imagine the zplane consists of two sheets superimposed on each other. We now cut the two sheets alongOBand join the lower edge of the bottom sheet to the upper edge of the top sheet. Then on starting in the bottom sheet and making one complete circuit about Owe arrive in the top sheet. We must now imagine the other cut edges to be joined together (independent of the first join and actually disregarding its existence) so that by continuing the circuit we go from the top sheet back to the bottom sheet. The collection of two sheets is called a Riemann surface corresponding to the function zp. Each sheet corresponds to a branch of the function and on each sheet the function is singled-valued. The concept of Riemann surfaces has the advantage that the various values of multiple-valued functions are obtained in acontinuous fashion. /C84he di/C128erential calculus of functions of a comple/C120 /C118ariable /C76imits and continuity The definitions of limits and continuity for functions of a complex variable are similar to those for a real variable. We say that f…z†has limit /C119 0aszapproaches 241Figure 6.5. Branch cut for the function /C119ˆ zp.DIFFERENTIAL CALCULUS z0, which is written as lim z!z0f…z†ˆ/C1190; …6:9† if (a)f…z†is defined and single-valued in a neighborhood of zˆz0, with the possible exception of the point z0itself; and (b) given any positive number /C34(however small), there exists a positive number such that f…z†ÿ/C1190 jj </C34whenever 0 <zÿz0 jj <. The limit must be independent of the manner in which zapproaches z0. Example 6.6 (a)I ff…z†ˆz2, prove that lim z!z0;f…z†ˆz2 0 (b) Find lim z!z0f…z†if f…z†ˆz2z6ˆz0 0 zˆz0:( Solution: (a) We must show that given any /C34/C620 we can find (depending in general on /C34) such that jz2ÿz2 0j</C34whenever 0 <jzÿz0j<. Now if 1, then 0 <jzÿz0j<implies that zÿz0 jj z‡z0 jj <z‡z0 jj ˆzÿz0‡2z0 jj ; z2ÿz2 0/C12/C12/C12/C12<…zÿz 0 jj ‡2z0jj † <1‡2z0jj …† : Taking as 1 or /C34=…1‡2jz0j†, whichever is smaller, we then have jz2ÿz2 0j</C34 whenever 0 <jzÿz0j<, and the required result is proved. (b) There is no di/C128erence between this problem and that in part (a), since in both cases we exclude zˆz0from consideration. Hence lim z!z0f…z†ˆz20. Note that the limit of f…z†asz!z0has nothing to do with the value of f…z†atz0. A function f…z†is said to be continuous at z0if, given any /C34/C620, there exists a /C620 such that f…z†ÿf…z0† jj </C34whenever 0 <zÿz0 jj <. This implies three conditions that must be met in order that f…z†be continuous at zˆz0: (1) lim z!z0f…z†ˆ/C1190must exist; (2)f…z0†must exist, that is, f…z†is defined at z0; (3)/C1190ˆf…z0†. For example, complex polynomials, 0‡ 1z1‡ 2z2‡ nzn(where imay be complex), are continuous everywhere. Quotients of polynomials are continuous whenever the denominator does not vanish. The following example provides further illustration. 242FUNCTIONS OF A COMPLE/C88 VARIABLE A function f…z†is said to be continuous in a region /C82of the zplane if it is continuous at all points of /C82. Points in the zplane where f…z†fails to be continuous are called discontinuities off…z†, and f…z†is said to be discontinuous at these points. If lim z!z0f…z†exists but is not equal to f…z0†, we call the point z0a removable discontinuity, since by redefining f…z0†to be the same as lim z!z0f…z†the function becomes continuous. To examine the continuity of f…z†atzˆ1 , we let zˆ1=/C119and examine the continuity of f…1=/C119†at/C119ˆ0. /C68erivatives and analytic functions Given a continuous, single-valued function of a complex variable f…z†in some region /C82of the zplane, the derivative f0…z†…df=dz†at some fixed point z0in/C82is defined as f0…z0†ˆlim z!0f…z0‡z†ÿf…z0† z; …6:10† provided the limit exists independently of the manner in which z!0. Here zˆzÿz0,a n d zis any point of some neighborhood of z0.I ff0…z†exists at z0 and every point zin some neighborhood of z0, then f…z†is said to be analytic at z0. And f…z†is analytic in a region /C82of the complex zplane if it is analytic at every point in /C82. In order to be analytic, f…z†must be single-valued and continuous. It is straightforward to see this. In view of Eq. (6.10), whenever f0…z0†exists, then lim z!0f…z0‡z†ÿf…z0† ‰Š ˆ lim z!0f…z0‡z†ÿf…z0† zlim z!0zˆ0 that is, lim z!0f…z†ˆf…z0†: Thus fis necessarily continuous at any point z0where its derivative exists. But the converse is not necessarily true, as the following example shows. Example 6.7 The function f…z†ˆz/C42 is continuous at z0, but dz/C42=dzdoes not exist anywhere. By definition, dz/C42 dzˆlim z!0…z‡z†/C42ÿz/C42 zˆlim x;y!0…x‡iy‡x‡iy†/C42ÿ…x‡iy†/C42 x‡iy ˆlim x;y!0xÿiy‡xÿiyÿ…xÿiy† x‡iyˆlim x;y!0xÿiy x‡iy: 243DIFFERENTIAL CALCULUS Ifyˆ0, the required limit is lim x!0x=xˆ1. On the other hand, if xˆ0, the required limit is ÿ1. Then since the limit depends on the manner in which z!0, the derivative does not exist and so f…z†ˆz/C42 is non-analytic everywhere. Example 6.8 Given f…z†ˆ2z2ÿ1, find f0…z†atz0ˆ1ÿi. Solution: f0…z0†ˆf0…1ÿi†ˆ lim z!1ÿi…2z2ÿ1†ÿ‰2…1ÿi†2ÿ1Š zÿ…1ÿi† ˆlim z!1ÿi2‰zÿ…1ÿi†Š‰z‡…1ÿi†Š zÿ…1ÿi† ˆlim z!1ÿi2‰z‡…1ÿi†Š ˆ4…1ÿi†: The rules for di/C128erentiating sums, products, and quotients are, in general, the same for complex functions as for real-valued functions. That is, if f0…z0†and /C1030…z0†exist, then: (1)…f‡/C103†0…z0†ˆf0…z0†‡/C1030…z0†; (2)…f/C103†0…z0†ˆf0…z0†/C103…z0†‡f…z0†/C1030…z0†; (3)f /C1030 …z0†ˆ/C103…z0†f0…z0†ÿf…z0†/C1030…z0† /C103…z0†2; if/C1030…z0†6 ˆ0: /C84he /C67auchy/C177Riemann conditions We call f…z†analytic at z0,i ff0…z†exists for all zin some neighborhood of z0; andf…z†is analytic in a region /C82if it is analytic at every point of /C82. Cauchy and Riemann provided us with a simple but extremely important test for the analyti- city of f…z†. To deduce the Cauchy–Riemann conditions for the analyticity of f…z†, let us return to Eq. (6.10): f0…z0†ˆlim z!0f…z0‡z†ÿf…z0† z: If we write f…z†ˆu…x;y†‡i/C118…x;y†, this becomes f0…z†ˆ lim x;y!0u…x‡x;y‡y†ÿu…x;y†‡i…same for /C118† x‡iy: There are of course an infinite number of ways to approach a point zon a two- dimensional surface. Let us consider two possible approaches – along xand along 244FUNCTIONS OF A COMPLE/C88 VARIABLE y. Suppose we first take the xroute, so yis fixed as we change x, that is, yˆ0 and x!0, and we have f0…z†ˆ lim x!0u…x‡x;y†ÿu…x;y† x‡i/C118…x‡x;y†ÿ/C118…x;y† x ˆ/C64u /C64x‡i/C64/C118 /C64x: We next take the yroute, and we have f0…z†ˆ lim y!0u…x;y‡y†ÿu…x;y† iy‡i/C118…x;y‡y†ÿ/C118…x;y† iy ˆÿi/C64u /C64y‡/C64/C118 /C64y: Now f…z†cannot possibly be analytic unless the two derivatives are identical. Thus a necessary condition for f…z†to be analytic is /C64u /C64x‡i/C64/C118 /C64xˆÿi/C64u /C64y‡/C64/C118 /C64y; from which we obtain /C64u /C64xˆ/C64/C118 /C64yand/C64u /C64yˆÿ/C64/C118 /C64x: …6:11† These are the Cauchy–Riemann conditions, named after the French mathemati- cian A. L. Cauchy (1789–1857) who discovered them, and the German mathema- tician Riemann who made them fundamental in his development of the theory ofanalytic functions. Thus if the function f…z†ˆu…x;y†‡i/C118…x;y†is analytic in a region /C82, then u…x;y†and/C118…x;y†satisfy the Cauchy–Riemann conditions at all points of /C82. Example 6.9Iff…z†ˆz 2ˆx2ÿy2‡2ixy, then f0…z†exists for all z/C58f0…z†ˆ2z,a n d /C64u /C64xˆ2xˆ/C64/C118 /C64y;and/C64u /C64yˆÿ2yˆÿ/C64/C118 /C64x: Thus, the Cauchy–Riemann equations (6.11) hold in this example at all points z. We can also find examples in which u…x;y†and/C118…x;y†satisfy the Cauchy– Riemann conditions (6.11) at zˆz0, but f0…z0†doesn’t exist. One such example is the following: f…z†ˆu…x;y†‡i/C118…x;y†ˆz5=jzj4ifz6ˆ0 0i f zˆ0( : The reader can show that u…x;y†and/C118…x;y†satisfy the Cauchy–Riemann condi- tions (6.11) at zˆ0, but that f0…0†does not exist. Thus f…z†is not analytic at zˆ0. The proof is straightforward, but very tedious. 245DIFFERENTIAL CALCULUS However, the Cauchy–Riemann conditions do imply analyticity provided an additional hypothesis is added: Given f…z†ˆu…x;y†‡i/C118…x;y†,i fu…x;y†and/C118…x;y†are contin- uous with continuous first partial derivatives and satisfy the Cauchy–Riemann conditions (11) at all points in a region /C82, then f…z†is analytic in /C82. To prove this, we need the following result from the calculus of real-valued functions of two variables: If /C104…x;y†;/C64/C104=/C64x, and /C64/C104=/C64yare continuous in some region /C82about …x0;y0†, then there exists a function H…x;y†such that H…x;y†!0a s…x;y†!… 0;0†and /C104…x0‡x;y0‡y†ÿ/C104…x0;y0†ˆ/C64/C104…x0;y0† /C64xx‡/C64/C104…x0;y0† /C64yy ‡H…x;y†  …x†2‡…y†2q : Let us return to lim z!0f…z0‡z†ÿf…z0† z; where z0is any point in region /C82and zˆx‡iy. Now we can write f…z0‡z†ÿf…z0†ˆ‰u…x0‡x;y0‡y†ÿu…x0;y0†Š ‡i‰/C118…x0‡x;y0‡y†ÿ/C118…x0;y0†Š ˆ/C64u…x0y0† /C64xx‡/C64u…x0y0† /C64yy‡H…x;y† …x† 2‡…y†2q ‡i/C64/C118…x0y0† /C64xx‡/C64/C118…x0y0† /C64yy ‡/C71…x;y† …x† 2‡…y†2q  ; where H…x;y†!0 and /C71…x;y†!0a s…x;y†!… 0;0†. Using the Cauchy–Riemann conditions and some algebraic manipulation we obtain f…z0‡z†ÿf…z0†ˆ/C64u…x0;y0† /C64x‡i/C64/C118…x0;y0† /C64x …x‡iy† ‡H…x;y†‡i/C71…x;y† ‰Š  …x†2‡…y†2q 246FUNCTIONS OF A COMPLE/C88 VARIABLE and f…z0‡z†ÿf…z0† zˆ/C64u…x0;y0† /C64x‡i/C64/C118…x0;y0† /C64x ‡H…xy†‡i/C71…xy† ‰Š  …x†2‡…y†2q x‡iy: But   …x†2‡…y†2q x‡iy/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ1: Thus, as z!0, we have …x;y†!… 0;0†and lim z!0f…z0‡z†ÿf…z0† zˆ/C64u…x0;y0† /C64x‡i/C64/C118…x0;y0† /C64x; which shows that the limit and so f0…z0†exist. Since f…z†is di/C128erentiable at all points in region /C82,f…z†is analytic at z0which is any point in /C82. The Cauchy–Riemann equations turn out to be both necessary and sucient conditions that f…z†ˆu…x;y†‡i/C118…x;y†be analytic. Analytic functions are also called regular or holomorphic functions. If f…z†is analytic everywhere in the finite zcomplex plane, it is called an entire function. A function f…z†is said to be singular at zˆz0, if it is not di/C128erentiable there; the point z0is called a singular point of f…z†. /C72armonic functions Iff…z†ˆu…x;y†‡i/C118…x;y†is analytic in some region of the zplane, then at every point of the region the Cauchy–Riemann conditions are satisfied: /C64u /C64xˆ/C64/C118 /C64y;and/C64u /C64yˆÿ/C64/C118 /C64x; and therefore /C642u /C64x2ˆ/C642/C118 /C64x/C64y;and/C642u /C64y2ˆÿ/C642/C118 /C64y/C64x; provided these second derivatives exist. In fact, one can show that if f…z†is analytic in some region /C82, all its derivatives exist and are continuous in /C82. Equating the two cross terms, we obtain /C642u /C64x2‡/C642u /C64y2ˆ0 …6:12a† throughout the region /C82. 247DIFFERENTIAL CALCULUS Similarly, by di/C128erentiating the first of the Cauch–Riemann equations with respect to y, the second with respect to x, and subtracting we obtain /C642/C118 /C64x2‡/C642/C118 /C64y2ˆ0: …6:12b† Eqs. (6.12a) and (6.12b) are Laplace’s partial di/C128erential equations in two inde- pendent variables xandy. Any function that has continuous partial derivatives of second order and that satisfies Laplace’s equation is called a harmonic function. We have shown that if f…z†ˆu…x;y†‡i/C118…x;y†is analytic, then both uand/C118are harmonic functions. They are called conjugate harmonic functions. This is a di/C128erent use of the word conjugate from that employed in determining z/C42. Given one of two conjugate harmonic functions, the Cauchy–Riemann equa- tions (6.11) can be used to find the other. Singular points A point at which f…z†fails to be analytic is called a singular point or a singularity off…z†; the Cauchy–Riemann conditions break down at a singularity. Various types of singular points exist. (1) Isolated singular points: The point zˆz0is called an isolated singular point off…z†if we can find /C620 such that the circle jzÿz0jˆencloses no singular point other than z0. If no such can be found, we call z0a non- isolated singularity. (2) Poles: If we can find a positive integer nsuch that limz!z0…zÿz0†nf…z†ˆA6ˆ0, then zˆz0is called a pole of order n.I f nˆ1,z0is called a simple pole. As an example, f…z†ˆ1=…zÿ2†has a simple pole at zˆ2. But f…z†ˆ1=…zÿ2†3has a pole of order 3 at zˆ2. (3) Branch point: A function has a branch point at z0if, upon encircling z0and returning to the starting point, the function does not return to the startingvalue. Thus the function is multiple-valued. An example is f…z†ˆ zp, which has a branch point at zˆ0. (4) Removable singularities: The singular point z 0is called a removable singu- larity of f…z†if lim z!z0f…z†exists. For example, the singular point at zˆ0 off…z†ˆsin…z†=zis a removable singularity, since lim z!0sin…z†=zˆ1. (5) Essential singularities: A function has an essential singularity at a point z0if it has poles of arbitrarily high order which cannot be eliminated by multi- plication by …zÿz0†n, which for any finite choice of n. An example is the function f…z†ˆe1=…zÿ2†, which has an essential singularity at zˆ2. (6) Singularities at infinity: The singularity of f…z†atzˆ1 is the same type as that of f…1=/C119†at/C119ˆ0. For example, f…z†ˆz2has a pole of order 2 at zˆ1 , since f…1=/C119†ˆ/C119ÿ2has a pole of order 2 at /C119ˆ0. 248FUNCTIONS OF A COMPLE/C88 VARIABLE /C69lementar/C121 functions of z /C84he exponential function ez/C40orexp(z)) The exponential function is of fundamental importance, not only for its own sake, but also as a basis for defining all the other elementary functions. In its definition we seek to preserve as many of the characteristic properties of the real exponential function exas possible. Specifically, we desire that: (a)ezis single-valued and analytic. (b)dez=dzˆez. (c)ezreduces to exwhen Im zˆ0: Recall that if we approach the point zalong the x-axis (that is, yˆ0; x!0), the derivative of an analytic function f0…z†can be written in the form f0…z†ˆdf dzˆ/C64u /C64x‡i/C64/C118 /C64x: If we let ezˆu‡i/C118; then to satisfy ( b) we must have /C64u /C64x‡i/C64/C118 /C64xˆu‡i/C118: Equating real and imaginary parts gives /C64u /C64xˆu; …6:13† /C64/C118 /C64xˆ/C118: …6:14† Eq. (6.13) will be satisfied if we write uˆex/C30…y†; …6:15† where /C30…y†is any function of y. Moreover, since ezis to be analytic, uand/C118must satisfy the Cauchy–Riemann equations (6.11). Then using the second of Eqs.(6.11), Eq. (6.14) becomes ÿ/C64u /C64yˆ/C118: 249ELEMENTARY FUNCTIONS OF z Di/C128erentiating this with respect to y, we obtain /C642u /C64y2ˆÿ/C64/C118 /C64y ˆÿ/C64u /C64x…with the aid of the f irst of Eqs :…6:11††: Finally, using Eq. (6.13), this becomes /C642u /C64y2ˆÿu; which, on substituting Eq. (6.15), becomes ex/C3000…y†ˆÿ ex/C30…y†or/C3000…y†ˆÿ /C30…y†: This is a simple linear di/C128erential equation whose solution is of the form /C30…y†ˆAcosy‡Bsiny: Then uˆex/C30…y† ˆex…Acosy‡Bsiny† and /C118ˆÿ/C64u /C64yˆÿex…ÿAsiny‡Bcosy†: Therefore ezˆu‡i/C118ˆex‰…Acosy‡Bsiny†‡i…AsinyÿBcosy†Š: If this is to reduce to exwhen yˆ0, according to ( c), we must have exˆex…AÿiB† from which we find Aˆ1 and Bˆ0: Finally we find ezˆex‡iyˆex…cosy‡isiny†: …6:16† This expression meets our requirements ( a), (b), and ( c); hence we adopt it as the definition of ez. It is analytic at each point in the entire zplane, so it is an entire function. Moreover, it satisfies the relation ez1ez2ˆez1‡z2: …6:17† It is important to note that the right hand side of Eq. (6.16) is in standard polar form with the modulus of ezgiven by exand an argument by y: mod ezjezjˆexand arg ezˆy: 250FUNCTIONS OF A COMPLE/C88 VARIABLE From Eq. (6.16) we obtain the Euler formula: eiyˆcosy‡isiny. Now let yˆ2, and since cos 2 ˆ1 and sin 2 ˆ0, the Euler formula gives e2iˆ1: Similarly, eiˆÿ1;ei=2ˆi: Combining this with Eq. (6.17), we find ez‡2iˆeze2iˆez; which shows that ezis periodic with the imaginary period 2 i. Thus ez2nIˆez…nˆ0;1;2;...†: …6:18† Because of the periodicity all the values that /C119ˆf…z†ˆezcan assume are already assumed in the strip ÿ<y. This infinite strip is called the fundamental region of ez. /C84rigonometric and hyperbolic functions From the Euler formula we obtain cosxˆ1 2…eix‡eÿix†;sinxˆ1 2i…eixÿeÿix†…xreal†: This suggests the following definitions for complex z: coszˆ12…e iz‡eÿiz†;sinzˆ1 2i…eizÿeÿiz†: …6:19† The other trigonometric functions are defined in the usual way: tanzˆsinz cosz;cotzˆcosz sinz;seczˆ1 cosz;cosec zˆ1 sinz; whenever the denominators are not zero. From these definitions it is easy to establish the validity of such familiar for- mulas as: sin…ÿz†ˆÿ sinz;cos…ÿz†ˆcosz;and cos2z‡sin2zˆ1; cos…z1z2†ˆcosz1cosz2/C7sinz1sinz2;sin…z1z2†ˆsinz1cosz2cosz1sinz2 dcosz dzˆÿsinz;dsinz dzˆcosz: Since ezis analytic for all z, the same is true for the function sin zand cos z. The functions tan zand sec zare analytic except at the points where cos zis zero, and cot zand cosec zare analytic except at the points where sin zis zero. The 251ELEMENTARY FUNCTIONS OF z functions cos zand sec zare even, and the other functions are odd. Since the exponential function is periodic, the trigonometric functions are also periodic, and we have cos…z2n†ˆcosz;sin…z2n†ˆsinz; tan…z2n†ˆtanz;cot…z2n†ˆcotz; where nˆ0;1;...: Another important property also carries over: sin zand cos zhave the same zeros as the corresponding real-valued functions: sinzˆ0 if and only if zˆn…ninteger †; coszˆ0 if and only if zˆ…2n‡1†=2…ninteger †: We can also write these functions in the form u…x;y†‡i/C118…x;y†. As an example, we give the details for cos z. From Eq. (6.19) we have coszˆ1 2…eiz‡eÿiz†ˆ12…e i…x‡iy†‡eÿi…x‡iy††ˆ12…e ÿyeix‡eyeÿix† ˆ12‰e ÿy…cosx‡isinx†‡ey…cosxÿisinx†Š ˆcosxey‡eÿy 2ÿisinxeyÿeÿy 2 or, using the definitions of the hyperbolic functions of real variables coszˆcos…x‡iy†ˆcosxcosh yÿisinxsinhy; similarly, sinzˆsin…x‡iy†ˆsinxcoshy‡icosxsinhy: In particular, taking xˆ0 in these last two formulas, we find cos…iy†ˆcosh y;sin…iy†ˆisinhy: There is a big di/C128erence between the complex and real sine and cosine func- tions. The real functions are bounded between ÿ1a n d ‡1, but the complex functions can take on arbitrarily large values. For example, if yis real, then cos iyˆ1 2…eÿy‡ey†!1 asy!1 ory!ÿ 1 . /C84he logarithmic function wˆlnz The real natural logarithm yˆlnxis defined as the inverse of the exponential function eyˆx. For the complex logarithm, we take the same approach and define /C119ˆlnzwhich is taken to mean that e/C119ˆz …6:20† for each z6ˆ0. 252FUNCTIONS OF A COMPLE/C88 VARIABLE Setting /C119ˆu‡i/C118andzˆreiˆjzjeiwe have e/C119ˆeu‡i/C118ˆeuei/C118ˆrei: It follows that euˆrˆzjjoruˆlnrˆlnzjj and /C118ˆˆargz: Therefore /C119ˆlnzˆlnr‡iˆlnzjj ‡ iargz: Since the argument of zis determined only in multiples of 2 , the complex natural logarithm is infinitely many-valued. If we let 1be the principal argument ofz, that is, the particular argument of zwhich lies in the interval 0 <2, then we can rewrite the last equation in the form lnzˆlnzjj‡i…‡2n† nˆ0;1;2;...: …6:21† For any particular value of n, a unique branch of the function is determined, and the logarithm becomes e/C128ectively single-valued. If nˆ0, the resulting branch of the logarithmic function is called the principal value. Any particular branch of the logarithmic function is analytic, for we have by di/C128erentiating the definitive rela- tionzˆe/C119, dz=d/C119ˆe/C119ˆzord/C119=dzˆd…lnz†=dzˆ1=z: For a particular value of nthe derivative of ln zthus exists for all z6ˆ0. For the real logarithm, yˆlnxmakes sense when x/C620. Now we can take a natural logarithm of a negative number, as shown in the following example. Example 6.10 lnÿ4ˆlnjÿ4j‡iarg…ÿ4†ˆln 4‡i…‡2n†; its principal value is ln 4 ‡i…†, a complex number. This explains why the logarithm of a negative number makesno sense in real variable. /C72yperbolic functions We conclude this section on ‘‘elementary functions’’ by mentioning briefly thehyperbolic functions; they are defined at points where the denominator does not vanish: sinhzˆ 1 2…ezÿeÿz†;coshzˆ12…e z‡eÿz†; tanh zˆsinhz=coshz;cothzˆcoshz=sinhz; sech zˆ1=cosh z;cosech zˆ1=sinhz: 253ELEMENTARY FUNCTIONS OF z Since ezandeÿzare entire functions, sinh zand cosh zare also entire functions. The singularities of tanh zand sech zoccur at the zeros of cosh z, and the singularities of coth zand cosech zoccur at the zeros of sinh z. As with the trigonometric functions, basic identities and derivative formulas carry over in the same form to the complex hyperbolic functions (just replace x byz). Hence we shall not list them here. /C67omple/C120 integration Complex integration is very important. For example, in applications we often en- counter real integrals which cannot be evaluated by the usual methods, but we canget help and relief from complex integration. In theory, the method of complex integration yields proofs of some basic properties of analytic functions, which would be very dicult to prove without using complex integration. The most fundamental result in complex integration is Cauchy’s integral theo- rem, from which the important Cauchy integral formula follows. These will be the subject of this section. /C76ine integrals in the complex plane As in real integrals, the indefinite integralRf…z†dzstands for any function whose derivative is f…z†. The definite integral of real calculus is now replaced by integrals of a complex function along a curve. Why/C63 To see this, we can express zin terms of a real parameter t:z…t†ˆx…t†‡iy…t†, where, say, atb. Now as tvaries from atob, the point ( x;y) describes a curve in the plane. We say this curve is smooth if there exists a tangent vector at all points on the curve; this means that dx/C47dt anddy/C47dt are continuous and do not vanish simultaneously for a<t<b. Let/C67be such a smooth curve in the complex zplane (Fig. 6.6), and we shall assume that /C67has a finite length (mathematicians call /C67a rectifiable curve). Let f…z†be continuous at all points of /C67. Subdivide /C67intonparts by means of points z 1;z2;...;znÿ1, chosen arbitrarily, and let aˆz0;bˆzn. On each arc joining zkÿ1 tozk(kˆ1;2;...;n) choose a point /C119k(possibly /C119kˆzkÿ1or/C119kˆzk) and form the sum SnˆXn kˆ1f…/C119k†zk zkˆzkÿzkÿ1: Now let the number of subdivisions nincrease in such a way that the largest of the chord lengths jzkjapproaches zero. Then the sum Snapproaches a limit. If this limit exists and has the same value no matter how the zjs and /C119js are chosen, then 254FUNCTIONS OF A COMPLE/C88 VARIABLE this limit is called the integral of f…z†along /C67and is denoted by Z Cf…z†dzorZb af…z†dz: …6:22† This is often called a contour integral (with contour /C67) or a line integral of f…z†. Some authors reserve the name contour integral for the special case in which /C67is a closed curve (so end aand end bcoincide), and denote it by the symbolH f…z†dz. We now state, without proof, a basic theorem regarding the existence of the contour integral: If /C67 is piecewise smooth and f(z) is continuous on /C67/C44 thenR Cf…z†dzexists . Iff…z†ˆu…x;y†‡i/C118…x;y†, the complex line integral can be expressed in terms of real line integrals as Z Cf…z†dzˆZ C…u‡i/C118†…dx‡idy†ˆZ C…udxÿ/C118dy†‡iZ C…/C118dx‡udy†;…6:23† where curve /C67may be open or closed but the direction of integration must be specified in either case. Reversing the direction of integration results in the change of sign of the integral. Complex integrals are, therefore, reducible to curvilinear real integrals and possess the following properties: (1)R C‰f…z†‡/C103…z†ŠdzˆR Cf…z†dz‡R C/C103…z†dz; (2)R Ckf…z†dzˆkR Cf…z†dz,kˆany constant (real or complex); (3)Rb af…z†dzˆÿRa bf…z†dz; (4)Rb af…z†dzˆRm af…z†d‡Rb mf…z†dz; (5)jR Cf…z†dzjML, where Mˆmaxjf…z†jon/C67, and Lis the length of /C67. Property (5) is very useful, because in working with complex line integrals it is often necessary to establish bounds on their absolute values. We now give a brief 255COMPLE/C88 INTEGRATION Figure 6.6. Complex line integral. proof. Let us go back to the definition: Z Cf…z†dzˆlim n!1Xn kˆ1f…/C119k†zk: Now Xn kˆ1f…/C119k†zk/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12X n kˆ1f…/C119k† jj zkjj MXn kˆ1zkjj ML; where we have used the fact that jf…z†j Mfor all points zon/C67and thatPjzkjrepresents the sum of all the chord lengths joining zkÿ1andzk, and that this sum is not greater than the length Lof/C67. Now taking the limit of both sides, and property (5) follows. It is possible to show, more generally, that Z Cf…z†dz/C12/C12/C12/C12/C12/C12/C12/C12Z Cf…z†jj dzjj: …6:24† Example 6.11 Evaluate the integralR C…z/C42†2dz, where /C67is a straight line joining the points zˆ0 andzˆ1‡2i. Solution: Since …z/C42†2ˆ…xÿiy†2ˆx2ÿy2ÿ2xyi; we have Z C…z/C42†2dzˆZ C‰…x2ÿy2†dx‡2xydyЇiZ C‰ÿ2xydx ‡…x2ÿy2†dyŠ: But the Cartesian equation of /C67isyˆ2x, and the above integral therefore becomes Z C…z/C42†2dzˆZ1 05x2dx‡iZ1 0…ÿ10x2†dxˆ5=3ÿi10=3: Example 6.12 Evaluate the integral Z Cdz …zÿz0†n‡1; where /C67is a circle of radius rand center at z0, and nis an integer. 256FUNCTIONS OF A COMPLE/C88 VARIABLE Solution: For convenience, let zÿz0ˆrei, where ranges from 0 to 2 asz ranges around the circle (Fig. 6.7). Then dzˆrieid, and the integral becomes Z2 0rieid rn‡1ei…n‡1†ˆi rnZ2 0eÿind: Ifnˆ0, this reduces to iZ2 0dˆ2i and if n6ˆ0, we have i rnZ2 0…cosnÿisinn†dˆ0: This is an important and useful result to which we will refer later. /C67auchy’s integral theorem Cauchy’s integral theorem has various theoretical and practical consequences. It states that if f…z†is analytic in a simply-connected region (domain) and on its boundary /C67, then I Cf…z†dzˆ0: …6:25† What do we mean by a simply-connected region/C63 A region /C82(mathematicians prefer the term ‘domain’) is called simply-connected if any simple closed curve which lies in /C82can be shrunk to a point without leaving /C82. That is, a simply- connected region has no hole in it (Fig. 6.7( a)); this is not true for a multiply- connected region. The multiply-connected regions of Fig. 6.7( b) and ( c)h a v e respectively one and three holes in them. 257COMPLE/C88 INTEGRATION Figure 6.7. Simply-connected and doubly-connected regions. Although a rigorous proof of Cauchy’s integral theorem is quite demanding and beyond the scope of this book, we shall sketch the main ideas. Note that the integral can be expressed in terms of two-dimensional vector fields /C65and/C66: I Cf…z†dzˆI C…udxÿ/C118dy†‡iZ C…/C118dx‡udy† ˆI C/C65…r†dr‡iI C/C66…r†dr; where /C65…r†ˆu^e1ÿ/C118^e2;/C66…r†ˆ/C118^e1‡u^e2: Applying Stokes’ theorem, we obtain I Cf…z†dzˆZZ Rda… /C114 /C65‡i/C114 /C66† ˆZZ Rdxdy ÿ/C64/C118 /C64x‡/C64u /C64y ‡i/C64u /C64xÿ/C64/C118 /C64y  ; where /C82is the region enclosed by /C67. Since f…x†satisfies the Cauchy–Riemann conditions, both the real and the imaginary parts of the integral are zero, thusproving Cauchy’s integral theorem. Cauchy’s theorem is also valid for multiply-connected regions. For simplicity we consider a doubly-connected region (Fig. 6.8). f…z†is analytic in and on the boundary of the region /C82between two simple closed curves C 1andC2. Construct a cross-cut AF. Then the region bounded by AB/C68EA/C70/C71/C72/C70A is simply-connected so by Cauchy’s theorem I Cf…z†dzˆI ABD/C69A/C70/C71H/C70Af…z†dzˆ0 or Z ABD/C69Af…z†dz‡Z A/C70f…z†dz‡Z /C70/C71H/C70f…z†dz‡Z /C70Af…z†dzˆ0: 258FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.8. Proof of Cauchy’s theorem for a doubly-connected region. ButR A/C70f…z†dzˆÿR /C70Af…z†dz, therefore this becomes Z ABD/C69Af…z†dzy‡Z /C70/C71H/C70f…z†dzyˆ0 or I Cf…z†dzˆI C1f…z†dz‡I C2f…z†dzˆ0; …6:26† where both C1andC2are traversed in the positive direction (in the sense that an observer walking on the boundary always has the region /C82on his left). Note that curves C1andC2are in opposite directions. If we reverse the direction of C2(now C2is also counterclockwise, that is, both C1andC2are in the same direction.), we have I C1f…z†dzÿI C2f…z†dzˆ0o rI C2f…z†dzˆI C1f…z†dz: Because of Cauchy’s theorem, an integration contour can be moved across any region of the complex plane over which the integrand is analytic without changing the value of the integral. It cannot be moved across a hole (the shaded area) or a singularity (the dot), but it can be made to collapse around one, as shown in Fig. 6.9. As a result, an integration contour /C67enclosing nholes or singularities can be replaced by nseparated closed contours Ci, each enclosing a hole or a singularity: I Cf…z†dzˆXn kˆ1I Cif…z†dz which is a generalization of Eq. (6.26) to multiply-connected regions. There is a converse of the Cauchy’s theorem, known as Morera’s theorem. We now state it without proof: Morera/C39s theorem: If f(z) is continuous in a simply/C45connected region /C82 and the /C67auchy/C39s theorem isvalid around every simple closed curve /C67 in /C82/C44 then f…z†is analytic in /C82. 259COMPLE/C88 INTEGRATION Figure 6.9. Collapsing a contour around a hole and a singularity. Example 6.13 EvaluateH Cdz=…zÿa†where /C67is any simple closed curve and zˆais (a) outside /C67,(b) inside /C67. Solution: (a)I fais outside /C67, then f…z†ˆ1=…zÿa†is analytic everywhere inside and on /C67. Hence by Cauchy’s theoremI Cdz=…zÿa†ˆ0: (b)I fais inside /C67and ÿis a circle of radius 2with center at zˆaso that ÿis inside C (Fig. 6.10). Then by Eq. (6.26) we haveI Cdz=…zÿa†ˆI ÿdz=…zÿa†: Now on ÿ,jzÿajˆ/C34,o rzÿaˆ/C34ei, then dzˆi/C34eid, and I ÿdz zÿaˆZ2 0i/C34eid /C34eiˆiZ2 0dˆ2i: /C67auchy’s integral formulas One of the most important consequences of Cauchy’s integral theorem is what isknown as Cauchy’s integral formula. It may be stated as follows. If f(z) is analytic in a simply/C45connected region /C82/C44 and z 0is any point in the interior of /C82 which is enclosed by a simple closed curve /C67/C44 then f…z0†ˆ1 2iI Cf…z† zÿz0dz; …6:27† the integration around /C67 being taken in the positive sense (counter/C45clockwise). 260FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.10. To prove this, let ÿbe a small circle with center at z0and radius r(Fig. 6.11), then by Eq. (6.26) we have I Cf…z† zÿz0dzˆI ÿf…z† zÿz0dz: Now jzÿz0jˆrorzÿz0ˆrei;0<2. Then dzˆireidand the integral on the right becomes I ÿf…z† zÿz0dzˆZ2 0f…z0‡rei†irei reidˆiZ2 0f…z0‡rei†d: Taking the limit of both sides and making use of the continuity of f…z†,w eh a v e I Cf…z† zÿz0dzˆlim r!0Z2 0f…z0‡rei†d ˆiZ2 0lim r!0f…z0‡rei†dˆiZ2 0f…z0†dˆ2if…z0†; from which we obtain f…z0†ˆ1 2iI Cf…z† zÿz0dzq:e:d: Cauchy’s integral formula is also true for multiply-connected regions, but we shall leave its proof as an exercise. It is useful to write Cauchy’s integral formula (6.27) in the form f…z†ˆ1 2iI Cf…z0†dz0 z0ÿz to emphasize the fact that zcan be any point inside the close curve /C67. Cauchy’s integral formula is very useful in evaluating integrals, as shown in the following example. 261COMPLE/C88 INTEGRATION Figure 6.11. Cauchy’s integral formula. Example 6.14 Evaluate the integralH Cezdz=…z2‡1†,i f/C67is a circle of unit radius with center at (a)zˆiand ( b)zˆÿi. Solution: (a) We first rewrite the integral in the form I Cez z‡idz zÿi; then we see that f…z†ˆez=…z‡i†and z0ˆi. Moreover, the function f…z†is analytic everywhere within and on the given circle of unit radius around zˆi. By Cauchy’s integral formula we have I Cez z‡idz zÿiˆ2if…i†ˆ2iei 2iˆ…cos 1‡isin 1†: (b) We find z0ˆÿiandf…z†ˆez=…zÿi†. Cauchy’s integral formula gives I Cez zÿidz z‡iˆÿ…cos 1ÿisin 1†: /C67auchy’s integral formula for higher derivatives Using Cauchy’s integral formula, we can show that an analytic function f…z†has derivatives of all orders given by the following formula: f…n†…z0†ˆn/C33 2iI Cf…z†dz …zÿz0†n‡1; …6:28† where /C67is any simple closed curve around z0andf…z†is analytic on and inside /C67. Note that this formula implies that each derivative of f…z†is itself analytic, since it possesses a derivative. We now prove the formula (6.28) by induction on n. That is, we first prove the formula for nˆ1: f0…z0†ˆ1 2iI Cf…z†dz …zÿz0†2: As shown in Fig. 6.12, both z0andz0‡/C104lie in /C82, and f0…z0†ˆlim /C104!0f…z0‡/C104†ÿf…z0† /C104: Using Cauchy’s integral formula we obtain 262FUNCTIONS OF A COMPLE/C88 VARIABLE f0…z0†ˆlim /C104!0f…z0‡/C104†ÿf…z0† /C104 ˆlim /C104!01 2i/C104I C1 zÿ…z0‡/C104†ÿ1 zÿz0/C26/C27 f…z†dz: Now 1 /C1041 zÿ…z0‡/C104†ÿ1 zÿz0 ˆ1 …zÿz0†2‡/C104 …zÿz0ÿ/C104†…zÿz0†2: Thus, f0…z0†ˆ1 2iI Cf…z† …zÿz0†2dz‡1 2ilim /C104!0/C104I Cf…z† …zÿz0ÿ/C104†…zÿz0†2dz: The proof follows if the limit on the right hand side approaches zero as /C104!0. To show this, let us draw a small circle ÿof radius centered at z0(Fig. 6.12), then 1 2ilim /C104!0/C104I Cf…z† …zÿz0ÿ/C104†…zÿz0†2dzˆ1 2ilim /C104!0/C104I ÿf…z† …zÿz0ÿ/C104†…zÿz0†2dz: Now choose hso small (in absolute value) that z0‡/C104lies in ÿandj/C104j< =2, and the equation for ÿisjzÿz0jˆ. Thus, we have jzÿz0ÿ/C104j jzÿz0jÿj/C104j/C62ÿ=2ˆ=2. Next, as f…z†is analytic in /C82, we can find a positive number Msuch that jf…z†j M. And the length of ÿis 2. Thus, /C104 2iI ÿf…z†dz …zÿz0ÿ/C104†…zÿz0†2/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 /C104jj 2M…2† …=2†…2†ˆ2/C104jjM 2!0a s /C104!0; proving the formula for f0…z0†. 263COMPLE/C88 INTEGRATION Figure 6.12. Fornˆ2, we begin with f0…z0‡/C104†ÿf0…z0† /C104ˆ1 2i/C104I C1 …zÿz0/C104†2ÿ1 …zÿz0†2() f…z†dz ˆ2/C33 2iI Cf…z† …zÿz0†3dz‡/C104 2iI C3…zÿz0†ÿ2/C104 …zÿz0ÿ/C104†2…zÿz0†3f…z†dz: The result follows on taking the limit as /C104!0 if the last term approaches zero. The proof is similar to that for the case nˆ1, for using the fact that the integral around /C67equals the integral around ÿ, we have /C104 2iI ÿ3…zÿz0†ÿ2/C104 …zÿz0ÿ/C104†2…zÿz0†3f…z†dz/C104jj 2M…2† …=2†23ˆ4/C104jjM 4;/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 assuming Mexists such that j‰3…zÿz 0†ÿ2/C104Šf…z†j<M. In a similar manner we can establish the results for nˆ3;4;.... We leave it to the reader to complete the proof by establishing the formula for f…n‡1†…z0†, assum- ing that f…n†…z0†is true. Sometimes Cauchy’s integral formula for higher derivatives can be used to evaluate integrals, as illustrated by the following example. Example 6.15 Evaluate I Ce2z …z‡1†4dz; where /C67is any simple closed path not passing through ÿ1. Consider two cases: (a)/C67does not enclose ÿ1. Then e2z=…z‡1†4is analytic on and inside /C67, and the integral is zero by Cauchy’s integral theorem. (b)/C67encloses ÿ1. Now Cauchy’s integral formula for higher derivatives applies. Solution: Letf…z†ˆe2z, then f…3†…ÿ1†ˆ3/C33 2iI Ce2z …z‡1†4dz: Now f…3†…ÿ1†ˆ8eÿ2, hence I Ce2z …z‡1†4dzˆ2i 3/C33f…3†…ÿ1†ˆ8 3eÿ2i: 264FUNCTIONS OF A COMPLE/C88 VARIABLE /C83eries representations of anal/C121tic functions We now turn to a very important notion: series representations of analytic func- tions. As a prelude we must discuss the notion of convergence of complex series. Most of the definitions and theorems relating to infinite series of real terms can be applied with little or no change to series whose terms are complex. /C67omplex sequences A complex sequence is an ordered list which assigns to each positive integer na complex number zn: z1;z2;...;zn;...: The numbers znare called the terms of the sequence. For example, both i;i2;...;in;...or 1‡i;…1‡i†=2;…1‡i†=4;…1‡i†=8;...are complex sequences. The nth term of the second sequence is (1 ‡i†=2nÿ1. A sequence z1;z2;...;zn;...is said to be convergent with the limit l(or simply to converge to the number l) if, given /C34/C620, we can find a positive integer /C78such that jznÿ/C108j</C34for each n/C78(Fig. 6.13). Then we write lim n!1znˆ/C108: In words, or geometrically, this means that each term znwith n/C62/C78(that is, z/C78;z/C78‡1;z/C78‡2;...†lies in the open circular region of radius /C34with center at l. In general, /C78depends on the choice of /C34. Here is an illustrative example. Example 6.17 Using the definition, show that lim n!1…1‡z=n†ˆ1 for all z. 265SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS Figure 6.13. Convergent complex sequence. Solution: Given any number /C34/C620, we must find /C78such that 1‡z nÿ1/C12/C12/C12/C12/C12/C12</C34 ; for all n/C62/C78 from which we find z=njj </C34 or zjj=n</C34 ifn/C62zjj=/C34/C78: Setting znˆxn‡iyn, we may consider a complex sequence z1;z2;...;znin terms of real sequences, the sequence of the real parts and the sequence of the imaginary parts: x1;x2;...;xn, and y1;y2;...;yn. If the sequence of the real parts converges to the number A, and the sequence of the imaginary parts converges to the number B, then the complex sequence z1;z2;...;znconverges to the limit A‡iB, as illustrated by the following example. Example 6.18Consider the complex sequence whose nth term is z nˆn2ÿ2n‡3 3n2ÿ4‡i2nÿ1 2n‡1: Setting znˆxn‡iyn, we find xnˆn2ÿ2n‡3 3n2ÿ4ˆ1ÿ…2=n†‡…3=n2† 3ÿ4=n2and ynˆ2nÿ1 2n‡1ˆ2ÿ1=n 2‡1=n: Asn!1 ;xn!1=3 and yn!1, thus, zn!1=3‡i. /C67omplex series We are interested in complex series whose terms are complex functions f1…z†‡f2…z†‡f3…z†‡‡ fn…z†‡ : …6:29† The sum of the first nterms is Sn…z†ˆf1…z†‡f2…z†‡f3…z†‡‡ fn…z†; which is called the nth partial sum of the series (6.29). The sum of the remaining terms after the nth term is called the remainder of the series. We can now associate with the series (6.29) the sequence of its partial sums S1;S2;...:If this sequence of partial sums is convergent, then the series converges; and if the sequence diverges, then the series diverges. We can put this in a formal way. The series (6.29) is said to converge to the sum S…z†in a region /C82if for any 266FUNCTIONS OF A COMPLE/C88 VARIABLE /C34/C620 there exists an integer /C78depending in general on /C34and on the particular value of zunder consideration such that Sn…z†ÿS…z† jj </C34 for all n/C62/C78 and we write lim n!1Sn…z†ˆS…z†: The di/C128erence Sn…z†ÿS…z†is just the remainder after nterms, Rn…z†; thus the definition of convergence requires that jRn…z†j ! 0a sn!1 . If the absolute values of the terms in (6.29) form a convergent series f1…z†jj ‡f2…z†jj ‡f3…z†jj ‡‡ fn…z†jj ‡ then series (6.29) is said to be absolutely convergent. If series (6.29) converges but is not absolutely convergent, it is said to be conditionally convergent. The terms of an absolutely convergent series can be rearranged in any manner whatsoever without a/C128ecting the sum of the series whereas rearranging the terms of a con- ditionally convergent series may alter the sum of the series or even cause the series to diverge. As with complex sequences, questions about complex series can also be reduced to questions about real series, the series of the real part and the series of theimaginary part. From the definition of convergence it is not dicult to prove the following theorem: A necessary and su/C129cient condition that the series of complex terms f 1…z†‡f2…z†‡f3…z†‡‡ fn…z†‡ should convergence is that the series of the real parts and the seriesof the imaginary parts of these terms should each converge. Moreover/C44 if X 1 nˆ1RefnandX1 nˆ1Imfn converge to the respective functions /C82(z) and I(z)/C44 then the given series converges to R…z†‡I…z†/C44 and the series f1…z†‡f2…z†‡f3…z†‡‡ fn…z†‡ converges to R…z†‡iI…z†. Of all the tests for the convergence of infinite series, the most useful is probably the familiar ratio test , which applies to real series as well as complex series. 267SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS Ratio test /C71iven the series f1…z†‡f2…z†‡f3…z†‡‡ fn…z†‡ /C44 the series converges abso/C45 lutely if 0<r…z†jj ˆ lim n!1fn‡1…z† fn…z†/C12/C12/C12/C12/C12/C12/C12/C12<1 …6:30† and diverges if jr…z†j/C621. /C87hen jr…z†j ˆ1/C44 the ratio test provides no information about the convergence or divergence of the series . Example 6.19 Consider the complex series X nSnˆX1 nˆ02ÿn‡ieÿn…† ˆX1 nˆ02ÿn‡iX1 nˆ0eÿn: The ratio tests on the real and imaginary parts show that both converge: lim n!12ÿ…n‡1† 2ÿn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ 1 2, which is positive and less than 1; lim n!1eÿ…n‡1† eÿn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ 1 e, which is also positive and less than 1. One can prove that the full series converges to X1 nˆ1Snˆ1 1ÿ1=2‡i1 1ÿeÿ1: /C85niform convergence and the /C87eierstrass M-test To establish conditions, under which series can legitimately be integrated or di/C128erentiated term by term, the concept of uniform convergence is required: A series of functions is said to converge uniformly to the functionS(z) in a region /C82/C44 either open or closed/C44 if corresponding to an arbitrary /C34<0there exists an integral /C78/C44 depending on /C34but not on z/C44 such that for every value of z in /C82 S…z†ÿS n…z† jj </C34 f/C111r a/C108/C108 n /C62/C78: One of the tests for uniform convergence is the Weierstrass M-test (a sucient test). 268FUNCTIONS OF A COMPLE/C88 VARIABLE If a se/C113uence of positive constants fMngexists such that jfn…z†j Mnfor all positive integers n and for all values of z in a given region /C82/C44 and if the series M1‡M2‡‡ Mn‡ is convergent/C44 then the series f1…z†‡f2…z†‡f3…z†‡‡ fn…z†‡ converges uniformly in /C82. As an illustrative example, we use it to test for uniform convergence of the series X1 nˆ1unˆX1 nˆ1zn n n‡1p in the region jzj1. Now junjˆjzjn n n‡1p 1 n3=2 ifjzj1. Calling Mnˆ1=n3=2, we see thatPMnconverges, as it is a pseries with /C112ˆ3=2. Hence by Wierstrass M-test the given series converges uniformly (and absolutely) in the indicated region jzj1. Po/C119er series and /C84aylor series Power series are one of the most important tools of complex analysis, as power series with non-zero radii of convergence represent analytic functions. As an example, the power series SˆX1 nˆ0anzn…6:31† clearly defines an analytic function as long as the series converge. We will only be interested in absolute convergence. Thus we have lim n!1an‡1zn‡1 anzn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12<1o r zjj<Rˆlim n!1anjj an‡1jj; where /C82is the radius of convergence since the series converges for all zlying strictly inside a circle of radius /C82centered at the origin. Similarly, the series SˆX1 nˆ0an…zÿz0†n converges within a circle of radius /C82centered at z0. 269SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS Notice that the Eq. (6.31) is just a Taylor series at the origin of a function with fn…0†ˆann/C33. Every choice we make for the infinite variables andefines a new function with its own set of derivatives at the origin. Of course we can go beyond the origin, and expand a function in a Taylor series centered at zˆz0. Thus in the complex analysis there is a Taylor expansion for every analytic function. This isthe question addressed by Taylor/C39s theorem (named after the English mathemati- cian Brook Taylor, 1685–1731): If f(z) is analytic throughout a region /C82 bounded by a simple closed curve /C67/C44 and if z and a are both interior to /C67/C44 then f(z) can be expanded in a Taylor series centered at zˆafor jzÿaj<R: f…z†ˆf…a†‡f 0…a†…zÿa†‡f00…a†…zÿa†2 2/C33‡  ‡fn…a†…zÿa†nÿ1 n/C33‡Rn; …6:32† where the remainder Rnis given by Rn…z†ˆ… zÿa†n1 2iI Cf…/C119†d/C119 …/C119ÿa†n…/C119ÿz†: Proof: To prove this, we first rewrite Cauchy’s integral formula as f…z†ˆ1 2iI Cf…/C119†d/C119 /C119ÿzˆ1 2iI Cf…/C119† /C119ÿa1 1ÿ…zÿa†=…/C119ÿa† d/C119:…6:33† For later use we note that since wis on /C67while zis inside /C67, zÿa /C119ÿa/C12/C12/C12/C12/C12/C12<1: From the geometric progression 1‡/C113‡/C113 2‡  ‡ /C113nˆ1ÿ/C113n‡1 1ÿ/C113ˆ1 1ÿ/C113ÿ/C113n‡1 1ÿ/C113 we obtain the relation 1 1ÿ/C113ˆ1‡/C113‡‡ /C113n‡/C113n‡1 1ÿ/C113: 270FUNCTIONS OF A COMPLE/C88 VARIABLE By setting /C113ˆ…zÿa†=…/C119ÿa†we find 1 1ÿ‰ …zÿa†=…/C119ÿa†Šˆ1‡zÿa /C119ÿa‡zÿa /C119ÿa2 ‡‡zÿa /C119ÿan ‡‰…zÿa†=…/C119ÿa†Šn‡1 …/C119ÿz†=…/C119ÿa†: We insert this into Eq. (6.33). Since zandaare constant, we may take the powers of (zÿa) out from under the integral sign, and then Eq. (6.33) takes the form f…z†ˆ1 2iI Cf…/C119†d/C119 /C119ÿa‡zÿa 2iI Cf…/C119†d/C119 …/C119ÿa†2‡‡…zÿa†n 2iI Cf…/C119†d/C119 …/C119ÿa†n‡1‡Rn…z†: Using Eq. (6.28), we may write this expansion in the form f…z†ˆf…a†‡zÿa 1/C33f0…a†‡…zÿa†2 2/C33f00…a†‡‡…zÿa†n n/C33fn…a†‡Rn…z†; where Rn…z†ˆ… zÿa†n1 2iI Cf…/C119†d/C119 …/C119ÿa†n…/C119ÿz†: Clearly, the expansion will converge and represent f…z†if and only if limn!1Rn…z†ˆ0. This is easy to prove. Note that wis on /C67while zis inside /C67,s ow eh a v e j/C119ÿzj/C620. Now f…z†is analytic inside /C67and on /C67, so it follows that the absolute value of f…/C119†=…/C119ÿz†is bounded, say, f…/C119† /C119ÿz/C12/C12/C12/C12/C12/C12/C12/C12<M for all won/C67. Let rbe the radius of /C67, then j/C119ÿajˆrfor all won/C67, and /C67has the length 2 r. Hence we obtain R njjˆjzÿajn 2I Cf…/C119†d/C119 …/C119ÿa†n…/C119ÿz†/C12/C12/C12/C12/C12/C12/C12/C12<zÿajjn 2M1 rn2r ˆMrzÿa r/C12/C12/C12/C12/C12/C12 n !0a s n!1 : Thus f…z†ˆf…a†‡zÿa 1/C33f0…a†‡…zÿa†2 2/C33f00…a†‡‡…zÿa†n n/C33fn…a† is a valid representation of f…z†at all points in the interior of any circle with its center at aand within which f…z†is analytic. This is called the Taylor series of f…z† with center at a. And the particular case where aˆ0 is called the Maclaurin series off…z†/C91Colin Maclaurin 1698–1746, Scots mathematician/C93. 271SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS The Taylor series of f…z†converges to f…z†only within a circular region around the point zˆa, the circle of convergence; and it diverges everywhere outside this circle. /C84aylor series of elementary functions Taylor series of analytic functions are quite similar to the familiar Taylor series of real functions. Replacing the real variable in the latter series by a complex vari- able we may ‘continue’ real functions analytically to the complex domain. The following is a list of Taylor series of elementary functions: in the case of multiple- valued functions, the principal branch is used. ezˆX1 nˆ0zn n/C33ˆ1‡z‡z2 2/C33‡ ; jzj<1; sinzˆX1 nˆ0…ÿ1†nz2n‡1 …2n‡1†/C33ˆzÿz3 3/C33‡z5 5/C33ÿ‡ ; jzj<1; coszˆX1 nˆ0…ÿ1†nz2n …2n†/C33ˆ1ÿz2 2/C33‡z4 4/C33ÿ‡ ; jzj<1; sinhzˆX1 nˆ0z2n‡1 …2n‡1†/C33ˆz‡z3 3/C33‡z5 5/C33‡ ; jzj<1; cosh zˆX1 nˆ0z2n …2n†/C33ˆ1‡z2 2/C33‡z4 4/C33‡ ; jzj<1; ln…1‡z†ˆX1 nˆ0…ÿ1†n‡1zn nˆzÿz2 2‡z3 3ÿ‡ ; jzj<1: Example 6.20Expand (1 ÿz† ÿ1about a. Solution: 1 1ÿzˆ1 …1ÿa†ÿ…zÿa†ˆ1 1ÿa1 1ÿ…zÿa†=…1ÿa†ˆ1 1ÿaX1 nˆ0zÿa 1ÿan : We have established two surprising properties of complex analytic functions: (1)They have derivatives of all order . (2)They can always be represented by Taylor series . This is not true in general for real functions; there are real functions which have derivatives of all orders but cannot be represented by a power series. 272FUNCTIONS OF A COMPLE/C88 VARIABLE Example 6.21 Expand ln( a‡z) about a. Solution: Suppose we know the Maclaurin series, then ln…1‡z†ˆln…1‡a‡zÿa†ˆln…1‡a†1‡zÿa 1‡a ˆln…1‡a†‡ln 1‡zÿa 1‡a ˆln…1‡a†‡zÿa 1‡a ÿ1 2zÿa 1‡a2 ‡13zÿa 1‡a3 ÿ‡ : Example 6.22 Letf…z†ˆln…1‡z†, and consider that branch which has the value zero when zˆ0. (a) Expand f…z†in a Taylor series about zˆ0, and determine the region of convergence. (b) Expand ln/C91(1 ‡z†=…1ÿz)/C93 in a Taylor series about zˆ0. Solution: (a) f…z†ˆln…1‡z† f…0†ˆ0 f0…z†ˆ… 1‡z†ÿ1f0…0†ˆ1 f00…z†ˆÿ … 1‡z†ÿ2f00…0†ˆÿ 1 fF…z†ˆ2…1‡z†ÿ3fF…0†ˆ2/C33 ...... f…n‡1†…z†ˆ… ÿ 1†nn/C33…1‡n†…n‡1†f…n‡1†…0†ˆ… ÿ 1†nn/C33: Then f…z†ˆln…1‡z†ˆf…0†‡f0…0†z‡f00…0† 2/C33z2‡fF…0† 3/C33z3‡ ˆzÿz2 2‡z3 3ÿz4 4‡ÿ : Thenth term is unˆ… ÿ 1†nÿ1zn=n. The ratio test gives lim n!1un‡1 un/C12/C12/C12/C12/C12/C12/C12/C12ˆlim n!1nz n‡1/C12/C12/C12/C12/C12/C12/C12/C12ˆzjj and the series converges for jzj<1. 273SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS (b)l n‰…1‡z†=…1ÿz†Š ˆln…1‡z†ÿln…1ÿz†. Next, replacing zbyÿzin Taylor’s expansion for ln …1‡z†, we have ln…1ÿz†ˆÿ zÿz2 2ÿz3 3ÿz4 4ÿ : Then by subtraction, we obtain ln1‡z 1ÿzˆ2z‡z3 3‡z5 5‡/C32! ˆX1 nˆ02z2n‡1 2n‡1: /C76aurent series In many applications it is necessary to expand a function f…z†around points where or in the neighborhood of which the function is not analytic. The Taylor series is not applicable in such cases. A new type of series known as the Laurentseries is required. The following is a representation which is valid in an annular ring bounded by two concentric circles of C 1andC2such that f…z†is single-valued and analytic in the annulus and at each point of C1andC2, see Fig. 6.14. The function f…z†may have singular points outside C1and inside C2. Hermann Laurent (1841–1908, French mathematician) proved that, at any point in theannular ring bounded by the circles, f…z†can be represented by the series f…z†ˆX 1 nˆÿ1an…zÿa†n…6:34† where anˆ1 2iI Cf…/C119†d/C119 …/C119ÿa†n‡1;nˆ0;1;2;...; …6:35† 274FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.14. Laurent theorem. each integral being taken in the counterclockwise sense around curve /C67lying in the annular ring and encircling its inner boundary (that is, /C67is any concentric circle between C1andC2). To prove this, let zbe an arbitrary point of the annular ring. Then by Cauchy’s integral formula we have f…z†ˆ1 2iI C1f…/C119†d/C119 /C119ÿz‡1 2iI C2f…/C119†d/C119 /C119ÿz; where C2is traversed in the counterclockwise direction and C2is traversed in the clockwise direction, in order that the entire integration is in the positive direction. Reversing the sign of the integral around C2and also changing the direction of integration from clockwise to counterclockwise, we obtain f…z†ˆ1 2iI C1f…/C119†d/C119 /C119ÿzÿ1 2iI C2f…/C119†d/C119 /C119ÿz: Now 1=…/C119ÿz†ˆ‰1=…/C119ÿa†Šf1=‰1ÿ…zÿa†=…/C119ÿa†Šg; ÿ1=…/C119ÿz†ˆ1=…zÿ/C119†ˆ‰1=…zÿa†Šf1=‰1ÿ…/C119ÿa†=…zÿa†Šg: Substituting these into f…z†we obtain: f…z†ˆ1 2iI C1f…/C119†d/C119 /C119ÿzÿ1 2iI C2f…/C119†d/C119 /C119ÿz ˆ1 2iI C1f…/C119† /C119ÿa1 1ÿ…zÿa†=…/C119ÿa† d/C119 ‡1 2iI C2f…/C119† zÿa1 1ÿ…/C119ÿa†=…zÿa† d/C119: Now in each of these integrals we apply the identity 1 1ÿ/C113ˆ1‡/C113‡/C1132‡  ‡ /C113nÿ1‡/C113n 1ÿ/C113 275SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS to the last factor. Then f…z†ˆ1 2iI C1f…/C119† /C119ÿa1‡zÿa /C119ÿa‡‡zÿa /C119ÿanÿ1 ‡…zÿa†n=…/C119ÿa†n 1ÿ…zÿa†=…/C119ÿa† d/C119 ‡1 2iI C2f…/C119† zÿa1‡/C119ÿa zÿa‡‡/C119ÿa zÿanÿ1 ‡…/C119ÿa†n=…zÿa†n 1ÿ…/C119ÿa†=…zÿa† d/C119 ˆ1 2iI C1f…/C119†d/C119 /C119ÿa‡zÿa 2iI C2f…/C119†d/C119 …/C119ÿa†2‡‡…zÿa†nÿ1 2iI C2f…/C119†d/C119 …/C119ÿa†n‡Rn1 ‡1 2i…zÿa†I C2f…/C119†d/C119‡1 2i…zÿa†2I C1…/C119ÿa†f…/C119†d/C119‡ ‡1 2i…zÿa†nI C1…/C119ÿa†nÿ1f…/C119†d/C119‡Rn2; where Rn1ˆ…zÿa†n 2iI C1f…/C119†d/C119 …/C119ÿa†n…/C119ÿz†; Rn2ˆ1 2i…zÿa†nI C2…/C119ÿa†nf…/C119†d/C119 zÿ/C119: The theorem will be established if we can show that lim n!1Rn2ˆ0 and lim n!1Rn1ˆ0. The proof of lim n!1Rn1ˆ0 has already been given in the derivation of the Taylor series. To prove the second limit, we note that for values ofwonC2 j/C119ÿajˆr1;jzÿajˆ/C26say;jzÿ/C119jˆj … zÿa†ÿ…/C119ÿa†j  /C26ÿr1; and j…f…/C119†j M; where Mis the maximum of jf…/C119†jonC2. Thus Rn2/C12/C12/C12/C12ˆ 1 2i…zÿa†nI C2…/C119ÿa†nf…/C119†d/C119 zÿ/C119/C12/C12/C12/C12/C12/C12/C12/C12 1 2ijj zÿajjnI C2/C119ÿajjnf…/C119†jj d/C119jj zÿ/C119jj or Rn2/C12/C12/C12/C12rn 1M 2/C26n…/C26ÿr1†I C2d/C119jjˆM 2r1 /C26n2r1 /C26ÿr1: 276FUNCTIONS OF A COMPLE/C88 VARIABLE Since r1=/C26 < 1, the last expression approaches zero as n!1 . Hence limn!1Rn2ˆ0 and we have f…z†ˆ1 2iI C1f…/C119†d/C119 /C119ÿa‡1 2iI C1f…/C119†d/C119 …/C119ÿa†2"# …zÿa† ‡1 2iI C1f…/C119†d/C119 …/C119ÿa†3"# …zÿa†2‡ ‡1 2iI C2f…/C119†d/C1191 zÿa‡1 2iI C2…/C119ÿa†f…/C119†d/C1191 …zÿa†2‡ : Since f…z†is analytic throughout the region between C1andC2, the paths of integration C1andC2can be replaced by any other curve /C67within this region and enclosing C2. And the resulting integrals are precisely the coecients angiven by Eq. (6.35). This proves the Laurent theorem. It should be noted that the coecients of the positive powers ( zÿa) in the Laurent expansion, while identical in form with the integrals of Eq. (6.28), cannot be replaced by the derivative expressions fn…a† n/C33 as they were in the derivation of Taylor series, since f…z†is not analytic through- out the entire interior of C2(or/C67), and hence Cauchy’s generalized integral formula cannot be applied. In many instances the Laurent expansion of a function is not found through the use of the formula (6.34), but rather by algebraic manipulations suggested by the nature of the function. In particular, in dealing with quotients of polynomials it isoften advantageous to express them in terms of partial fractions and then expand the various denominators in series of the appropriate form through the use of the binomial expansion, which we assume the reader is familiar with: …s‡t† nˆsn‡nsnÿ1tn…nÿ1† 2/C33snÿ2t2‡n…nÿ1†…nÿ2† 3/C33snÿ3t3‡ : This expansion is valid for all values of n if jsj/C62jtj:Ifjsjjtjthe expansion is valid only if n is a non/C45negative integer. That such procedures are correct follows from the fact that the Laurent expan/C45 sion of a function over a given annular ring is uni/C113ue . That is, if an expansion of the Laurent type is found by any process, it must be the Laurent expansion. Example 6.23 Find the Laurent expansion of the function f…z†ˆ… 7zÿ2†=‰…z‡1†z…zÿ2†Šin the annulus 1 <jz‡1j<3. 277SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS Solution: We first apply the method of partial fractions to f…z†and obtain f…z†ˆÿ3 z‡1‡1 z‡2 zÿ2: Now the center of the given annulus is zˆÿ1, so the series we are seeking must be one involving powers of z‡1. This means that we have to modify the second and third terms in the partial fraction representation of f…z†: f…z†ˆÿ3 z‡1‡1 …z‡1†ÿ1‡2 …z‡1†ÿ3; but the series for ‰…z‡1†ÿ3Šÿ1converges only where jz‡1j/C623, whereas we require an expansion valid for jz‡1j<3. Hence we rewrite the third term in the other order: f…z†ˆÿ3 z‡1‡1 …z‡1†ÿ1‡2 ÿ3‡…z‡1† ˆÿ3…z‡1†ÿ1‡‰ …z‡1†ÿ1Šÿ1‡2‰ÿ3‡…z‡1†Šÿ1 ˆ‡… z‡1†ÿ2ÿ2…z‡1†ÿ1ÿ2 3ÿ29…z‡1† ÿ2 27…z‡1†2ÿ ; 1<jz‡1j<3: Example 6.24 Given the following two functions: …a†e3z…z‡1†ÿ3; …b†…z‡2†sin1 z‡2; find Laurent series about the singularity for each of the functions, name the singularity, and give the region of convergence. Solution: (a)zˆÿ1 is a triple pole (pole of order 3). Let z‡1ˆu, then zˆuÿ1 and e3z …z‡1†3ˆe3…uÿ1† u3ˆeÿ3e3u u3ˆeÿ3 u31‡3u‡…3u†2 2/C33‡…3u†3 3/C33‡…3u†4 4/C33‡/C32! ˆeÿ3 1 …z‡1†3‡3 …z‡1†2‡9 2…z‡1†‡9 2‡27…z‡1† 8‡/C32! : The series converges for all values of z6ˆ ÿ1. (b)zˆÿ2 is an essential singularity. Let z‡2ˆu, then zˆuÿ2, and 278FUNCTIONS OF A COMPLE/C88 VARIABLE …z‡2†sin1 z‡2ˆusin1 uˆu1uÿ1 3/C33u3‡1 5/C33u5‡ ˆ1ÿ1 6…z‡2†2‡1 120…z‡2†4ÿ‡ : The series converges for all values of z6ˆ ÿ2. Integration b/C121 the method of residues We now turn to integration by the method of residues which is useful in evaluat- ing both real and complex integrals. We first discuss briefly the theory of residues, then apply it to evaluate certain types of real definite integrals occurring in physics and engineering. Residues Iff…z†is single-valued and analytic in a neighborhood of a point zˆa, then, by Cauchy’s integral theorem, I Cf…z†dzˆ0 for any contour in that neighborhood. But if f…z†has a pole or an isolated essential singularity at zˆaand lies in the interior of /C67, then the above integral will, in general, be di/C128erent from zero. In this case we may represent f…z†by a Laurent series: f…z†ˆX1 nˆÿ1an…zÿa†nˆa0‡a1…zÿa†‡a2…zÿa†2‡‡aÿ1 zÿa‡aÿ2 …zÿa†2‡ ; where anˆ1 2iI Cf…z† …zÿa†n‡1dz; nˆ0;1;2;...: The sum of all the terms containing negative powers, namely aÿ1=…zÿa†‡aÿ2=…zÿa†2‡ ;is called the principal part of f…z†atzˆa.I n the special case nˆÿ1, we have aÿ1ˆ1 2iI Cf…z†dz or I Cf…z†dzˆ2iaÿ1; …6:36† 279INTEGRATION BY THE METHOD OF RESIDUES the integration being taken in the counterclockwise sense around a simple closed curve /C67that lies in the region 0 <jzÿaj<Dand contains the point zˆa, where /C68is the distance from ato the nearest singular point of f…z†. The coecient aÿ1is called the residue of f…z†atzˆa, and we shall use the notation aÿ1ˆRes zˆaf…z†: …6:37† We have seen that Laurent expansions can be obtained by various methods, without using the integral formulas for the coecients. Hence, we may determine the residue by one of those methods and then use the formula (6.36) to evaluate contour integrals. To illustrate this, let us consider the following simple example. Example 6.25 Integrate the function f…z†ˆzÿ4sinzaround the unit circle /C67in the counter- clockwise sense. Solution: Using sinzˆX1 nˆ0…ÿ1†nz2n‡1 …2n‡1†/C33ˆzÿz3 3/C33‡z5 5/C33ÿ‡ ; we obtain the Laurent series f…z†ˆsinz z4ˆ1 z3ÿ1 3/C33z‡z 5/C33ÿz3 7/C33‡ÿ : We see that f…z†has a pole of third order at zˆ0, the corresponding residue is aÿ1ˆÿ1=3/C33, and from Eq. (6.36) it follows that Isinz z4dzˆ2iaÿ1ˆÿi 3: There is a simple standard method for determining the residue in the case of a pole. If f…z†has a simple pole at a point zˆa, the corresponding Laurent series is of the form f…z†ˆX1 nˆÿ1an…zÿa†nˆa0‡a1…zÿa†‡a2…zÿa†2‡‡aÿ1 zÿa; where aÿ16ˆ0. Multiplying both sides by zÿa,w eh a v e …zÿa†f…z†ˆ… zÿa†‰a0‡a1…zÿa†‡ Ї aÿ1 and from this we have Res zˆaf…z†ˆaÿ1ˆlim z!a…zÿa†f…z†: …6:38† 280FUNCTIONS OF A COMPLE/C88 VARIABLE Another useful formula is obtained as follows. If f…z†can be put in the form f…z†ˆ/C112…z† /C113…z†; where /C112…z†and/C113…z†are analytic at zˆa;/C112…z†6 ˆ0, and /C113…z†ˆ0a tzˆa(that is, /C113…z†has a simple zero at zˆa†. Consequently, /C113…z†can be expanded in a Taylor series of the form /C113…z†ˆ… zÿa†/C1130…a†‡…zÿa†2 2/C33/C11300…a†‡ : Hence Res zˆaf…z†ˆlim z!a…zÿa†/C112…z† /C113…z†ˆlim z!a…zÿa†/C112…z† …zÿa†‰/C1130…a†‡…zÿa†/C11300…a†=2‡ Šˆ/C112…a† /C1130…a†: …6:39† Example 6.26 The function f…z†ˆ… 4ÿ3z†=…z2ÿz†is analytic except at zˆ0 and zˆ1 where it has simple poles. Find the residues at these poles. Solution: We have p…z†ˆ4ÿ3z;/C113…z†ˆz2ÿz. Then from Eq. (6.39) we obtain Res zˆ0f…z†ˆ4ÿ3z 2zÿ1 zˆ0ˆÿ4; Res zˆ1f…z†ˆ4ÿ3z 2zÿ1 zˆ1ˆ1: We now consider poles of higher orders. If f…z†has a pole of order m/C621a ta point zˆa, the corresponding Laurent series is of the form f…z†ˆa0‡a1…zÿa†‡a2…zÿa†2‡‡aÿ1 zÿa‡aÿ2 …zÿa†2‡‡aÿm …zÿa†m; where aÿm6ˆ0 and the series converges in some neighborhood of zˆa, except at the point itself. By multiplying both sides by …zÿa†mwe obtain …zÿa†mf…z†ˆaÿm‡aÿm‡1…zÿa†‡aÿm‡2…zÿa†2‡‡ aÿm‡…mÿ1†…zÿa†…mÿ1† ‡…zÿa†m‰a0‡a1…zÿa†‡ Š : This represents the Taylor series about zˆaof the analytic function on the left hand side. Di/C128erentiating both sides ( mÿ1) times with respect to z,w eh a v e dmÿ1 dzmÿ1‰…zÿa†mf…z†Š ˆ … mÿ1†/C33aÿ1‡m…mÿ1†2a0…zÿa†‡ : 281INTEGRATION BY THE METHOD OF RESIDUES Thus on letting z!a lim z!admÿ1 dzmÿ1‰…zÿa†mf…z†Š ˆ … mÿ1†/C33aÿ1; that is, Res zˆaf…z†ˆ1 …mÿ1†/C33lim z!admÿ1 dzmÿ1…zÿa†mf…z† ‰Š() : …6:40† Of course, in the case of a rational function f…z†the residues can also be determined from the representation of f…z†in terms of partial fractions. /C84he residue theorem So far we have employed the residue method to evaluate contour integrals whose integrands have only a single singularity inside the contour of integration. Now consider a simple closed curve /C67containing in its interior a number of isolated singularities of a function f…z†. If around each singular point we draw a circle so small that it encloses no other singular points (Fig. 6.15), these small circles, together with the curve /C67, form the boundary of a multiply-connected region in which f…z†is everywhere analytic and to which Cauchy’s theorem can therefore be applied. This gives 1 2iI Cf…z†dz‡I C1f…z†dz‡‡I Cmf…z†dz ˆ0: If we reverse the direction of integration around each of the circles and change thesign of each integral to compensate, this can be written 1 2iI Cf…z†dzˆ1 2iI C1f…z†dz‡1 2iI C2f…z†dz‡‡1 2iI Cmf…z†dz; 282FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.15. Residue theorem. where all the integrals are now to be taken in the counterclockwise sense. But the integrals on the right are, by definition, just the residues of f…z†at the various isolated singularities within /C67. Hence we have established an important theorem, the residue theorem: Iff…z†is analytic inside a simple closed curve /C67 and on /C67 ,except at a /C174nite number of singular points a1/C44a2;...;amin the interior of C/C44 then I Cf…z†dzˆ2iXm jˆ1Res zˆajf…z†ˆ2i…r1‡r2‡‡ rm†;…6:41† where rjis the residue of f…z†at the singular point aj. Example 6.27 The function f…z†ˆ… 4ÿ3z†=…z2ÿz†has simple poles at zˆ0 and zˆ1; the residues are ÿ4 and 1, respectively (cf. Example 6.26). ThereforeI C4ÿ3z z2ÿzdzˆ2i…ÿ4‡1†ˆÿ 6i for every simple closed curve /C67which encloses the points 0 and 1, andI C4ÿ3z z2ÿzdzˆ2i…ÿ4†ˆÿ 8i for any simple closed curve /C67for which zˆ0 lies inside /C67andzˆ1 lies outside, the integrations being taken in the counterclockwise sense. /C69/C118aluation of real definite integrals The residue theorem yields a simple and elegant method for evaluating certainclasses of complicated real definite integrals. One serious restriction is that the contour must be closed. But many integrals of practical interest involve integra- tion over open curves. Their paths of integration must be closed before the residue theorem can be applied. So our ability to evaluate such an integral depends crucially on how the contour is closed, since it requires knowledge of the addi- tional contributions from the added parts of the closed contour. A number oftechniques are known for closing open contours. The following types are most common in practice. Improper integrals of the rational functionZ 1 ÿ1f…x†dx The improper integral has the meaning Z1 ÿ1f…x†dxˆlim a!1Z0 af…x†dx‡lim b!1Zb 0f…x†dx: …6:42† 283EVALUATION OF REAL DEFINITE INTEGRALS If both limits exist, we may couple the two independent passages to ÿ1 and1, and write Z1 ÿ1f…x†dxˆlim r!1Zr ÿrf…x†dx: …6:43† We assume that the function f…x†is a real rational function whose denominator is di/C128erent from zero for all real xand is of degree at least two units higher than the degree of the numerator. Then the limits in (6.42) exist and we can start from (6.43). We consider the corresponding contour integral I Cf…z†dz; along a contour /C67consisting of the line along the x-axis from ÿrtorand the semicircle ÿabove (or below) the x-axis having this line as its diameter (Fig. 6.16). Then let r!1 .I ff…x†is an even function this can be used to evaluate Z1 0f…x†dx: Let us see why this works. Since f…x†is rational, f…z†has finitely many poles in the upper half-plane, and if we choose rlarge enough, /C67encloses all these poles. Then by the residue theorem we have I Cf…z†dzˆZ ÿf…z†dz‡Zr ÿrf…x†dxˆ2iX Resf…z†: This gives Zr ÿrf…x†dxˆ2iX Resf…z†ÿZ ÿf…z†dz: We next prove thatR ÿf…z†dz!0i fr!1 . To this end, we set zˆrei, then ÿ is represented by rˆconst, and as zranges along ÿ;ranges from 0 to . Since 284FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.16. Path of the contour integral. the degree of the denominator of f…z†is at least 2 units higher than the degree of the numerator, we have f…z†jj <k=zjj2…zjjˆr/C62r0† for suciently large constants kandr. By applying (6.24) we thus obtain Z ÿf…z†dz/C12/C12/C12/C12/C12/C12/C12/C12<k r2rˆk r: Hence, as r!1 , the value of the integral over ÿapproaches zero, and we obtain Z1 ÿ1f…x†dxˆ2iX Resf…z†: …6:44† Example 6.28 Using (6.44), show that Z1 0dx 1‡x4ˆ 2 2p: Solution: f …z†ˆ1=…1‡z4†has four simple poles at the points z1ˆei=4;z2ˆe3i=4;z3ˆeÿ3i=4;z4ˆeÿi=4: The first two poles, z1andz2, lie in the upper half-plane (Fig. 6.17) and we find, using L’Hospital’s rule Res zˆz1f…z†ˆ1 …1‡z4†0 zˆz1ˆ1 4z3 zˆz1ˆ1 4eÿ3i=4ˆÿ14e i=4; Res zˆz2f…z†ˆ1 …1‡z4†0 zˆz2ˆ1 4z3 zˆz2ˆ14e ÿ9i=4ˆ14e ÿi=4; 285EVALUATION OF REAL DEFINITE INTEGRALS Figure 6.17. then Z1 ÿ1dx 1‡x4ˆ2i 4…ÿei=4‡eÿi=4†ˆsin 4ˆ 2p and so Z1 0dx 1‡x4ˆ1 2Z1 ÿ1dx 1‡x4ˆ 2 2p: Example 6.29 Show that Z1 ÿ1x2dx …x2‡1†2…x2‡2x‡2†ˆ7 50: Solution: The poles of f…z†ˆz2 …z2‡1†2…z2‡2z‡2† enclosed by the contour of Fig. 6.17 are zˆiof order 2 and zˆÿ1‡iof order 1. The residue at zˆiis lim z!id dz…zÿi†2 z2 …z‡i†21…zÿi†2…z2‡2z‡2†"# ˆ9iÿ12 100: The residue at zˆÿ1‡iis lim z!ÿ1‡i…z‡1ÿi†z2 …z2‡1†2…z‡1ÿi†…z‡1‡i†ˆ3ÿ4i 25: Therefore Z1 ÿ1x2dx …x2‡1†2…x2‡2x‡2†ˆ2i9iÿ12 100‡3ÿ4i 25 ˆ7 50: Integrals of the rational functions of sinandcosZ2 0/C71…sin;cos†d /C71…sin;cos†is a real rational function of sin and cos finite on the interval 02. Let zˆei, then dzˆieid;ordˆdz=iz;sinˆ…zÿzÿ1†=2i;cosˆ…z‡zÿ1†=2 286FUNCTIONS OF A COMPLE/C88 VARIABLE and the given integrand becomes a rational function of z, say, f…z†.A s ranges from 0 to 2 , the variable zranges once around the unit circle jzjˆ1 in the counterclockwise sense. The given integral takes the formI Cf…z†dz iz; the integration being taken in the counterclockwise sense around the unit circle. Example 6.30 Evaluate Z2 0d 3ÿ2c o s ‡sin: Solution: Letzˆei, then dzˆieid,o rdˆdz=iz, and sinˆzÿzÿ1 2i;cosˆz‡zÿ1 2; then Z2 0d 3ÿ2 cos ‡sinˆI C2dz …1ÿ2i†z2‡6izÿ1ÿ2i; where /C67is the circle of unit radius with its center at the origin (Fig. 6.18). We need to find the poles of 1 …1ÿ2i†z2‡6izÿ1ÿ2i/C58 zˆÿ6i …6i†2ÿ4…1ÿ2i†…ÿ1ÿ2i†q 2…1ÿ2i† ˆ2ÿi;…2ÿi†=5; 287EVALUATION OF REAL DEFINITE INTEGRALS Figure 6.18. only (2 ÿi†=5 lies inside /C67, and residue at this pole is lim z!…2ÿi†=5‰zÿ…2ÿi†=5Š2 …1ÿ2i†z2‡6izÿ1ÿ2i ˆ lim z!…2ÿi†=52 2…1ÿ2i†z‡6iˆ1 2iby L’Hospital’s rule : Then Z2 0d 3ÿ2c o s ‡sinˆI C2dz …1ÿ2i†z2‡6izÿ1ÿ2iˆ2i…1=2i†ˆ: Fourier integrals of the formZ1 ÿ1f…x†sinmx cosmx/C26/C27 dx Iff…x†is a rational function satisfying the assumptions stated in connection with improper integrals of rational functions, then the above integrals may be evalu- ated in a similar way. Here we consider the corresponding integral I Cf…z†eimzdz over the contour /C67as that in improper integrals of rational functions (Fig. 6.16), and obtain the formula Z1 ÿ1f…x†eimxdxˆ2iX Res‰f…z†eimzŠ…m/C620†; …6:45† where the sum consists of the residues of f…z†eimzat its poles in the upper half- plane. Equating the real and imaginary parts on each side of Eq. (6.45), we obtain Z1 ÿ1f…x†cosmxdx ˆÿ2X Im Res ‰f…z†eimzŠ; …6:46† Z1 ÿ1f…x†sinmxdx ˆ2X Re Res ‰f…z†eimzŠ: …6:47† To establish Eq. (6.45) we should now prove that the value of the integral over the semicircle ÿin Fig. 6.16 approaches zero as r!1 . This can be done as follows. Since ÿlies in the upper half-plane y0a n d m/C620, it follows that jeimzjˆjeimxjeÿmyjj ˆeÿmy1 …y0;m/C620†: From this we obtain jf…z†eimzjˆf…z†jj jeimzjf…z†jj …y0;m/C620†; which reduces our present problem to that of an improper integral of a rational function of this section, since f…x†is a rational function satisfying the assumptions 288FUNCTIONS OF A COMPLE/C88 VARIABLE stated in connection these improper integrals. Continuing as before, we see that the value of the integral under consideration approaches zero as rapproaches 1, and Eq. (6.45) is established. Example 6.31 Show that Z1 ÿ1cosmx k2‡x2dxˆ keÿkm;Z1 ÿ1sinmx k2‡x2dxˆ0 …m/C620;k/C620†: Solution: The function f…z†ˆeimz=…k2‡z2†has a simple pole at zˆikwhich lies in the upper half-plane. The residue of f…z†atzˆikis Res zˆikeimz k2‡z2ˆeimz 2z zˆikˆeÿmk 2ik: Therefore Z1 ÿ1eimx k2‡x2dxˆ2ieÿmk 2ikˆ keÿmk and this yields the above results. Other types of real improper integrals These are definite integrals ZB Af…x†dx whose integrand becomes infinite at a point ain the interval of integration, limx!af…x†jj ˆ1 . This means that ZB Af…x†dxˆlim /C34!0Zaÿ/C34 Af…x†dx‡lim /C17!0Z a‡/C17f…x†dx; where both /C34and/C17approach zero independently and through positive values. It may happen that neither of these limits exists when /C34; /C17!0 independently, but lim /C34!0Zaÿ/C34 Af…x†dx‡ZB a‡/C34f…x†dx exists; this is called Cauchy’s principal value of the integral and is often written pr:v:ZB Af…x†dx: 289EVALUATION OF REAL DEFINITE INTEGRALS To evaluate improper integrals whose integrands have poles on the real axis, we can use a path which avoids these singularities by following small semicircles with centers at the singular points. We now illustrate the procedure with a simple example. Example 6.32 Show that Z1 0sinx xdxˆ 2: Solution: The function sin …z†=zdoes not behave suitably at infinity. So we con- sider eiz=z, which has a simple pole at zˆ0, and integrate around the contour /C67 orAB/C68E/C70/C71A (Fig. 6.19). Since eiz=zis analytic inside and on /C67, it follows from Cauchy’s integral theorem that I Ceiz zdzˆ0 or Zÿ/C34 ÿReix xdx‡Z C2eiz zdz‡ZR /C34eix xdx‡Z C1eiz zdzˆ0: …6:48† We now prove that the value of the integral over large semicircle C1approaches zero as /C82approaches infinity. Setting zˆRei, we have dzˆiReid;dz=zˆid and therefore Z C1eiz zdz/C12/C12/C12/C12/C12/C12/C12/C12ˆZ  0eizid/C12/C12/C12/C12/C12/C12/C12/C12Z  0eiz/C12/C12/C12/C12d: In the integrand on the right, eiz/C12/C12/C12/C12ˆje iR…cos‡isin†jˆjeiRcosjjeÿRsinjˆeÿRsin: 290FUNCTIONS OF A COMPLE/C88 VARIABLE Figure 6.19. By inserting this and using sin( ÿ†ˆsinwe obtain Z 0eiz/C12/C12/C12/C12dˆZ  0eÿRsindˆ2Z=2 0eÿRsind ˆ2Z/C34 0eÿRsind‡Z=2 /C34eÿRsind"# ; where /C34has any value between 0 and =2. The absolute values of the integrands in the first and the last integrals on the right are at most equal to 1 and eÿRsin/C34, respectively, because the integrands are monotone decreasing functions of in the interval of integration. Consequently, the whole expression on the right is smaller than 2Z/C34 0d‡eÿRsinZ=2 /C34d ˆ2/C34‡eÿRsin 2ÿ/C34  <2/C34‡eÿRsin/C34: Altogether Z C1eiz zdz/C12/C12/C12/C12/C12/C12/C12/C12<2/C34‡eÿRsin/C34: We first take /C34arbitrarily small. Then, having fixed /C34, the last term can be made as small as we please by choosing /C82suciently large. Hence the value of the integral along C1approaches 0 as R!1 . We next prove that the value of the integral over the small semicircle C2 approaches zero as /C34!0. Let zˆ/C34i, then Z C2eiz zdzˆÿlim /C34!0Z0 exp…i/C34ei† /C34eii/C34eidˆÿlim /C34!0Z0 iexp…i/C34ei†dˆi and Eq. (6.48) reduces to Zÿ/C34 ÿReix xdx‡i‡ZR /C34eix xdxˆ0: Replacing xbyÿxin the first integral and combining with the last integral, we find ZR /C34eixÿeÿix xdx‡iˆ0: Thus we have 2iZR /C34sinx xdxˆi: 291EVALUATION OF REAL DEFINITE INTEGRALS Taking the limits R!1 and/C34!0 Z1 0sinx xdxˆ 2: Problems 6.1. Given three complex numbers z1ˆa‡ib,z2ˆc‡id, and z3ˆ/C103‡i/C104, show that: (a)z1‡z2ˆz2‡z1 commutative law of addition; (b)z1‡…z2‡z3†ˆ… z1‡z2†‡z3 associative law of addition; (c)z1z2ˆz2z1 commutative law of multiplication; (d)z1…z2z3†ˆ… z1z2†z3 associative law of multiplication. 6.2. Given z1ˆ3‡4i 3ÿ4i;z2ˆ1‡2i 1ÿ3i2 find their polar forms, complex conjugates, moduli, product, the quotient z1=z2: 6.3. The absolute value or modulus of a complex number zˆx‡iyis defined as zjjˆ zz/C42p ˆ x2‡y2q : Ifz1;z2;...;zmare complex numbers, show that the following hold: (a)jz1z2jˆjz1jjz2jorjz1z2zmjˆjz1jjz2jjzmj: (b)jz1=z2jˆjz1j=jz2jifz26ˆ0: (c)jz1‡z2jjz1j‡jz2j: (d)jz1‡z2jjz1jÿjz2jorjz1ÿz2jjz1jÿjz2j. 6.4 Find all roots of ( a) ÿ325p , and ( b) 1‡i3p , and locate them in the complex plane. 6.5 Show, using De Moivre’s theorem, that: (a) cos 5 ˆ16 cos5ÿ20 cos3‡5c o s ; (b) sin 5 ˆ5 cos4sinÿ10 cos2sin3‡sin5. 6.6 Given zˆrei, interpret zei, where is real geometrically. 6.7 Solve the quadratic equation az2‡bz‡cˆ0;a6ˆ0. 6.8 A point Pmoves in a counterclockwise direction around a circle of radius 1 with center at the origin in the zplane. If the mapping function is /C119ˆz2, show that when Pmakes one complete revolution the image P0ofPin the w plane makes three complete revolutions in a counterclockwise direction on a circle of radius 1 with center at the origin. 6.9 Show that f…z†ˆlnzhas a branch point at zˆ0. 6.10 Let /C119ˆf…z†ˆ… z2‡1†1=2, show that: 292FUNCTIONS OF A COMPLE/C88 VARIABLE (a)f…z†has branch points at zˆI. (b) a complete circuit around both branch points produces no change in the branches of f…z†. 6.11 Apply the definition of limits to prove that: lim z!1z2ÿ1 zÿ1ˆ2: 6.12. Prove that: (a)f…z†ˆz2is continuous at zˆz0,a n d (b)f…z†ˆz2;z6ˆz0 0;zˆz0( is discontinuous at zˆz0, where z06ˆ0. 6.13 Given f…z†ˆz/C42, show that f0…i†does not exist. 6.14 Using the definition, find the derivative of f…z†ˆz3ÿ2zat the point where: (a)zˆz0, and (b) zˆÿ1. 6.15. Show that fis an analytic function of zif it does not depend on z/C42/C58f…z;z/C42†ˆf…z†. In other words, f…x;y†ˆf…x‡iy†, that is, xand y enter fonly in the combination x/C43iy . 6.16. (a) Show that uˆy3ÿ3x2yis harmonic. (b) Find /C118such that f…z†ˆu‡i/C118is analytic. 6.17 ( a)I ff…z†ˆu…x;y†‡i/C118…x;y†is analytic in some region /C82of the zplane, show that the one-parameter families of curves u…x;y†ˆC1and /C118…x;y†ˆC2are orthogonal families. (b) Illustrate ( a) by using f…z†ˆz2. 6.18 For each of the following functions locate and name the singularities in the finite zplane: (a)f…z†ˆz …z2‡4†4;(b)f…z†ˆsin zp  zp ;(c)f…z†ˆP1 nˆ01 znn/C33: 6.19 ( a) Locate and name all the singularities of f…z†ˆz8‡z4‡2 …zÿ1†3…3z‡2†2: (b) Determine where f…z†is analytic. 6.20 ( a) Given ezˆex…cosy‡isiny†, show that …d=dz†ezˆez. (b) Show that ez1ez2ˆez1‡z2. (Hint: set z1ˆx1‡iy1andz2ˆx2‡iy2and apply the addition formulas for the sine and cosine.) 6.21 Show that: ( a)l nezˆz‡2ni,(b)l nz1=z2ˆlnz1ÿlnz2‡2ni. 6.22 Find the values of: (a) ln i,(b)l n ( 1 ÿi). 6.23 EvaluateR Cz/C42dzfrom zˆ0t ozˆ4‡2ialong the curve /C67given by: (a)zˆt2‡it; (b) the line from zˆ0t ozˆ2iand then the line from zˆ2itozˆ4‡2i. 293PROBLEMS 6.24 EvaluateH Cdz=…zÿa†n;nˆ2;3;4;...where zˆais inside the simple closed curve /C67. 6.25 If f…z†is analytic in a simply-connected region /C82, and aandzare any two points in /C82, show that the integral Zz af…z†dz is independent of the path in /C82joining aandz. 6.26 Let f…z†be continuous in a simply-connected region /C82and let aandzbe points in /C82. Prove that /C70…z†ˆRz af…z0†dz0is analytic in /C82, and /C700…z†ˆf…z†. 6.27 Evaluate (a)I Csinz2‡cosz2 …zÿ1†…zÿ2†dz (b)I Ce2z …z‡1†4dz, where /C67is the circle jzjˆ1. 6.28 Evaluate I C2 sinz2 …zÿ1†4dz; where /C67is any simple closed path not passing through 1. 6.29 Show that the complex sequence znˆ1 nÿn2ÿ1 ni diverges. 6.30 Find the region of convergence of the seriesP1 nˆ1…z‡2†n‡1=…n‡1†34n. 6.31 Find the Maclaurin series of f…z†ˆ1=…1‡z2†. 6.32 Find the Taylor series of f…z†ˆsinzabout zˆ=4, and determine its circle of convergence. (Hint: sin zˆsin‰a‡…zÿa†Š:† 6.33 Find the Laurent series about the indicated singularity for each of the following functions. Name the singularity in each case and give the region of convergence of each series. (a)…zÿ3†sin1 z‡2;zˆÿ2; (b)z …z‡1†…z‡2†;zˆÿ2; (c)1 z…zÿ3†2;zˆ3: 6.34 Expand f…z†ˆ1=‰…z‡1†…z‡3†Šin a Laurent series valid for: (a)1<jzj<3, ( b)jzj/C623, ( c)0<jz‡1j<2. 294FUNCTIONS OF A COMPLE/C88 VARIABLE 6.35 Evaluate Z1 ÿ1x2dx …x2‡a2†…x2‡b2†; a/C620;b/C620: 6.36 Evaluate …a†Z2 0d 1ÿ2/C112cos‡/C1122; where pis a fixed number in the interval 0 </C112<1; …b†Z2 0d …5ÿ3 sin †2: 6.37 Evaluate Z1 ÿ1xsinx x2‡2x‡5dx: 6.38 Show that: …a†Z1 0sinx2dxˆZ1 0cosx2dxˆ1 2 2/C114 ; …b†Z1 0x/C112ÿ1 1‡xdxˆ sin/C112; 0</C112<1: 295PROBLEMS 7 Special functions of mathematical physics The functions discussed in this chapter arise as solutions of second-order di/C128er- ential equations which appear in special, rather than in general, physical pro- blems. So these functions are usually known as the special functions of mathematical physics. We start with Legendre’s equation (Adrien MarieLegendre, 1752–1833, French mathematician). Legendre/C39s equation Legendre’s di/C128erential equation …1ÿx 2†d2y dx2ÿ2xdy dx‡/C23…/C23‡1†yˆ0; …7:1† where vis a positive constant, is of great importance in classical and quantum physics. The reader will see this equation in the study of central force motion inquantum mechanics. In general, Legendre’s equation appears in problems inclassical mechanics, electromagnetic theory, heat, and quantum mechanics, with spherical symmetry. Dividing Eq. (7.1) by 1 ÿx 2, we obtain the standard form d2y dx2ÿ2x 1ÿx2dy dx‡/C23…/C23‡1† 1ÿx2yˆ0: We see that the coecients of the resulting equation are analytic at xˆ0, so the origin is an ordinary point and we may write the series solution in the form yˆX1 mˆ0amxm: …7:2† 296 Substituting this and its derivatives into Eq. (7.1) and denoting the constant /C23…/C23‡1†bykwe obtain …1ÿx2†X1 mˆ2m…mÿ1†amxmÿ2ÿ2xX1 mˆ1mamxmÿ1‡kX1 mˆ0amxmˆ0: By writing the first term as two separate series we have X1 mˆ2m…mÿ1†amxmÿ2ÿX1 mˆ2m…mÿ1†amxmÿ2X1 mˆ1mamxm‡kX1 mˆ0amxmˆ0; which can be written as: 21a2‡32a3x‡43a4x2‡‡… s‡2†…s‡1†as‡2xs‡ ÿ21a2x2ÿ ÿ… s…sÿ1†asxsÿ ÿ21a1xÿ22a2x2ÿ ÿ 2sasxsÿ ‡ka0‡ka1x ‡ka2x2‡ ‡ kasxs‡ˆ 0: Since this must be an identity in xif Eq. (7.2) is to be a solution of Eq. (7.1), the sum of the coecients of each power of xmust be zero; remembering that kˆ/C23…/C23‡1†we thus have 2a2‡/C23…/C23‡1†a0ˆ0; …7:3a† 6a3‡‰ ÿ 2‡/C118…/C118‡1†Ša1ˆ0; …7:3b† and in general, when sˆ2;3;...; …s‡2†…s‡1†as‡2‡‰ ÿs…sÿ1†ÿ2s‡/C23…/C23‡1†Šasˆ0: …4:4† The expression in square brackets /C91 .../C93 can be written …/C23ÿs†…/C23‡s‡1†: We thus obtain from Eq. (7.4) as‡2ˆÿ…/C23ÿs†…/C23‡s‡1† …s‡2†…s‡1†as…sˆ0;1;...†: …7:5† This is a recursion formula, giving each coecient in terms of the one two places before it in the series, except for a0anda1, which are left as arbitrary constants. 297LEGENDRE’S EQUATION We find successively a2ˆÿ/C23…/C23‡1† 2/C33a0; a3ˆÿ…/C23ÿ1†…/C23‡2† 3/C33a1; a4ˆÿ…/C23ÿ2†…/C23‡3† 43a2; a5ˆÿ…/C23ÿ3†…/C23‡4† 3/C33a3; ˆ…/C23ÿ2†/C23…/C23‡1†…/C23‡3† 4/C33a0; ˆ…/C23ÿ3†…/C23ÿ1†…/C23‡2†…/C23‡4† 5/C33a1; etc. By inserting these values for the coecients into Eq. (7.2) we obtain y…x†ˆa0y1…x†‡a1y2…x†; …7:6† where y1…x†ˆ1ÿ/C23…/C23‡1† 2/C33x2‡…/C23ÿ2†/C23…/C23‡1†…/C23‡3† 4/C33x4ÿ‡ … 7:7a† and y2…x†ˆxˆ…/C23ÿ1†…/C23‡2† 3/C33x3‡…/C23ÿ2†…/C23ÿ1†…/C23‡2†…/C23‡4† 5/C33x5ÿ‡ :…7:7b† These series converge for jxj<1. Since Eq. (7.7a) contains even powers of x, and Eq. (7.7b) contains odd powers of x, the ratio y1=y2is not a constant, and y1and y2are linearly independent solutions. Hence Eq. (7.6) is a general solution of Eq. (7.1) on the interval ÿ1<x<1. In many applications the parameter /C23in Legendre’s equation is a positive integer n. Then the right hand side of Eq. (7.5) is zero when sˆnand, therefore, an‡2ˆ0 and an‡4ˆ0;...:Hence, if nis even, y1…x†reduces to a polynomial of degree n.I fnis odd, the same is true with respect to y2…x†. These polynomials, multiplied by some constants, are called Legendre polynomials. Since they are of great practical importance, we will consider them in some detail. For this purpose we rewrite Eq. (7.5) in the form asˆÿ…s‡2†…s‡1† …nÿs†…n‡s‡1†as‡2 …7:8† and then express all the non-vanishing coecients in terms of the coecient anof the highest power of xof the polynomial. The coecient anis then arbitrary. It is customary to choose anˆ1 when nˆ0 and anˆ…2n†/C33 2n…n/C33†2ˆ135…2nÿ1† n/C33; nˆ1;2;...; …7:9† 298SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS the reason being that for this choice of anall those polynomials will have the value 1 when xˆ1. We then obtain from Eqs. (7.8) and (7.9) anÿ2ˆÿn…nÿ1† 2…2nÿ1†anˆÿn…nÿ1†…2n†/C33 2…2nÿ1†2n…n/C33†2 ˆÿn…nÿ1†2n…2nÿ1†…2nÿ2†/C33/C33 2…2nÿ1†2nn…nÿ1†/C33n…nÿ1†…nÿ2†/C33; that is, anÿ2ˆÿ…2nÿ2†/C33 2n…nÿ1†/C33…nÿ2†/C33: Similarly, anÿ4ˆÿ…nÿ2†…nÿ3† 4…2nÿ3†anÿ2ˆ…2nÿ4†/C33 2n2/C33…nÿ2†/C33…nÿ4†/C33 etc., and in general anÿ2mˆ… ÿ 1†m …2nÿ2m†/C33 2nm/C33…nÿm†/C33…nÿ2m†/C33: …7:10† The resulting solution of Legendre’s equation is called the Legendre polynomial of degree nand is denoted by Pn…x†; from Eq. (7.10) we obtain Pn…x†ˆXM mˆ0…ÿ1†m …2nÿ2m†/C33 2nm/C33…nÿm†/C33…nÿ2m†/C33xnÿ2m ˆ…2n†/C33 2n…n/C33†2xnÿ…2nÿ2†/C33 2n1/C33…nÿ1†/C33…nÿ2†/C33xnÿ2‡ÿ ;…7:11† where Mˆn=2o r…nÿ1†=2, whichever is an integer. In particular (Fig. 7.1) P0…x†ˆ1;P1…x†ˆx;P2…x†ˆ1 2…3x2ÿ1†;P3…x†ˆ12…5x3ÿ3x†; P4…x†ˆ18…35x4ÿ30x2‡3†;P5…x†ˆ18…63x5ÿ70x3‡15x†: Rodrigues’ formula for Pn…x† The Legendre polynomials Pn…x†are given by the formula Pn…x†ˆ1 2nn/C33dn dxn‰…x2ÿ1†nŠ: …7:12† We shall establish this result by actually carrying out the indicated di/C128erentia- tions, using the Leibnitz rule for nth derivative of a product, which we state below without proof: 299LEGENDRE’S EQUATION If we write DnuasunandDn/C118as/C118n, then …u/C118†nˆu/C118n‡nC1u1/C118nÿ1‡‡nCrur/C118nÿr‡‡ un/C118; where Dˆd=dxandnCris the binomial coecient and is equal to n/C33=‰r/C33…nÿr†/C33Š. We first notice that Eq. (7.12) holds for nˆ0, 1. Then, write zˆ…x2ÿ1†n=2nn/C33 so that …x2ÿ1†Dzˆ2nxz: …7:13† Di/C128erentiating Eq. (7.13) …n‡1†times by the Leibnitz rule, we get …1ÿx2†Dn‡2zÿ2xDn‡1z‡n…n‡1†Dnzˆ0: Writing yˆDnz, we then have: (i)yis a polynomial. (ii) The coecient of xnin…x2ÿ1†nis…ÿ1†n=2nCn=2(neven) or 0 ( nodd). Therefore the lowest power of xiny…x†isx0(neven) or x1(nodd). It follows that yn…0†ˆ0 …nodd† and yn…0†ˆ1 2nn/C33…ÿ1†n=2nCn=2n/C33ˆ…ÿ1†n=2n/C33 2n‰…n=2†/C33Š2…neven†: 300SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Figure 7.1. Legendre polynomials. By Eq. (7.11) it follows that yn…0†ˆPn…0†…alln†: (iii)…1ÿx2†D2yÿ2xDy‡n…n‡1†yˆ0, which is Legendre’s equation. Hence Eq. (7.12) is true for all n. /C84he generating function for Pn…x† One can prove that the polynomials Pn…x†are the coecients of znin the expan- sion of the function …x;z†ˆ… 1ÿ2xz‡z2†ÿ1=2, with jzj<1; that is, …x;z†ˆ… 1ÿ2xz‡z2†ÿ1=2ˆX1 nˆ0Pn…x†zn; zjj<1: …7:14† …x;z†is called the generating function for Legendre polynomials Pn…x†. We shall be concerned only with the case in which xˆcos…ÿ< † and then z2ÿ2xz‡1…zÿei†…zÿei†: The expansion (7.14) is therefore possible when jzj<1. To prove expansion (7.14) we have lhsˆ1‡1 2z…2xÿ1†‡13 222/C33z2…2xÿz†2‡ ‡13…2nÿ1† 2nn/C33zn…2xÿz†n‡ : The coecient of znin this power series is 13…2nÿ1† 2nn/C33…2nxn†‡13…2nÿ3† 2nÿ1…nÿ1†/C33‰ÿ…nÿ1†…2x†nÿ2Їˆ Pn…x† by Eq. (7.11). We can use Eq. (7.14) to find successive polynomials explicitly. Thus, di/C128erentiating Eq. (7.14) with respect to zso that …xÿz†…1ÿ2xz‡z2†ÿ3=2ˆX1 nˆ1nznÿ1Pn…x† and using Eq. (7.14) again gives …xÿz†P0…x†‡X1 nˆ1Pn…x†zn"# ˆ…1ÿ2xz‡z2†X1 nˆ1nznÿ1Pn…x†: …7:15† 301LEGENDRE’S EQUATION Then expanding coecients of znin Eq. (7.15) leads to the recurrence relation …2n‡1†xPn…x†ˆ…n‡1†Pn‡1…x†‡nPnÿ1…x†: …7:16† This gives P4;P5;P6, etc. very quickly in terms of P0;P1, and P3. Recurrence relations are very useful in simplifying work, helping in proofs or derivations. We list four more recurrence relations below without proofs or derivations: xP0 n…x†ÿP0 nÿ1…x†ˆnPn…x†; …7:16a† P0 n…x†ÿxP0 nÿ1…x†ˆnPnÿ1…x†; …7:16b† …1ÿx2†P0 n…x†ˆnPnÿ1…x†ÿnxP n…x†; …7:16c† …2n‡1†Pn…x†ˆP0 n‡1…x†ÿP0 nÿ1…x†: …7:16d† With the help of the recurrence formulas (7.16) and (7.16b), it is straight- forward to establish the other three. Omitting the full details, which are left for the reader, these relations can be obtained as follows: (i) di/C128erentiation of Eq. (7.16) with respect to xand the use of Eq. (7.16b) to eliminate P0 n‡1…x†leads to relation (7.16a); (ii) the addition of Eqs. (7.16a) and (7.16b) immediately yields relation (7.16d); (iii) the elimination of P0 nÿ1…x†between Eqs. (7.16b) and (7.16a) gives relation (7.16c). Example 7.1The physical significance of expansion (7.14) is apparent in this simple example: find the potential /C86of a point charge at point Pdue to a charge ‡/C113at/C81. Solution: Suppose the origin is at O(Fig. 7.2). Then /C86 Pˆ/C113 Rˆ/C113…/C262ÿ2r/C26cos‡r2†ÿ1=2: 302SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Figure 7.2. Thus, if r</C26 /C86Pˆ/C113 /C26…1ÿ2zcos‡z2†ÿ1=2; zˆr=/C26; which gives /C86Pˆ/C113 /C26X1 nˆ0r /C26n Pn…cos†… r</C26†: Similarly, when r/C62/C26, we get /C86Pˆ/C113 /C26X1 nˆ0/C26 rn‡1 Pn…cos†: There are many problems in which it is essential that the Legendre polynomials be expressed in terms of , the colatitude angle of the spherical coordinate system. This can be done by replacing xby cos . But this will lead to expressions that are quite inconvenient because of the powers of cos they contain. Fortunately, using the generating function provided by Eq. (7.14), we can derive more useful forms in which cosines of multiples of take the place of powers of cos . To do this, let us substitute xˆcosˆ…ei‡eÿi†=2 into the generating function, which gives ‰1ÿz…ei‡eÿi†‡z2Šÿ1=2ˆ‰ …1ÿzei†…1ÿzeÿi†Šÿ1=2ˆX1 nˆ0Pn…cos†zn: Now by the binomial theorem, we have …1ÿzei†ÿ1=2ˆX1 nˆ0anzneni; …1ÿzeÿi†ÿ1=2ˆX1 nˆ0anzneÿni; where anˆ135…2nÿ1† 246…2n†; n1; a0ˆ1: …7:17† To find the coecient of znin the product of these two series, we need to form the Cauchy product of these two series. What is a Cauchy product of two series/C63 We state it below for the reader who is in need of a review: The Cauchy product of two infinite series,P1 nˆ0un…x† andP1 nˆ0/C118n…x†, is defined as the sum over n X1 nˆ0sn…x†ˆX1 nˆ0Xn kˆ0uk…x†/C118nÿk…x†; 303LEGENDRE’S EQUATION where sn…x†is given by sn…x†ˆXn kˆ0uk…x†/C118nÿk…x†ˆu0…x†/C118n…x†‡‡ un…x†/C1180…x†: Now the Cauchy product for our two series is given by X1 nˆ0Xn kˆ0anÿkznÿke…nÿk†i akzkeÿki ˆX1 nˆ0 znXn kˆ0akanÿke…nÿ2k†i :…7:18† In the inner sum, which is the sum of interest to us, it is straightforward to prove that, for n1, the terms corresponding to kˆjandkˆnÿjare identical except that the exponents on eare of opposite sign. Hence these terms can be paired, and we have for the coecient of zn, Pn…cos†ˆa0an…eni‡eÿni†‡a1anÿ1…e…nÿ2†i‡eÿ…nÿ2†i†‡ ˆ2a0ancosn‡a1anÿ1cos…nÿ2†‡ ‰Š :…7:19† Ifnis odd, the number of terms is even and each has a place in one of the pairs. In this case, the last term in the sum is a…nÿ1†=2a…n‡1†=2cos: Ifnis even, the number of terms is odd and the middle term is unpaired. In this case, the series (7.19) for Pn…cos†ends with the constant term an=2an=2: Using Eq. (7.17) to compute values of the an, we find from the unit coecient of z0 in Eqs. (7.18) and (7.19), whether nis odd or even, the specific expressions P0…cos†ˆ1; P1…cos†ˆcos;P2…cos†ˆ… 3 cos 2 ‡1†=4 P3…cos†ˆ… 5 cos 3 ‡3 cos †=8 P4…cos†ˆ… 35 cos 4 ‡20 cos 2 ‡9†=64 P5…cos†ˆ… 63 cos 5 ‡35 cos 3 ‡30 cos †=128 P6…cos†ˆ… 231 cos 6 ‡126 cos 4 ‡105 cos 2 ‡50†=5129 >>>>>>>>= >>>>>>>>;:…7:20† Orthogonality of /C76egendre polynomials The set of Legendre polynomials fP n…x†gis orthogonal for ÿ1x‡1. In particular we can show that Z‡1 ÿ1Pn…x†Pm…x†dxˆ2=…2n‡1†ifmˆn 0i f m6ˆn:/C26 …7:21† 304SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS (i)m6ˆn: Let us rewrite the Legendre equation (7.1) for Pm…x†in the form d dx…1ÿx2†P0 m…x†/C2/C3 ‡m…m‡1†Pm…x†ˆ0 …7:22† and the one for Pn…x† d dx…1ÿx2†P0 n…x†/C2/C3 ‡n…n‡1†Pn…x†ˆ0: …7:23† We then multiply Eq. (7.22) by Pn…x†and Eq. (7.23) by Pm…x†, and subtract to get Pmd dx…1ÿx2†P0 n/C2/C3 ÿPnd dx…1ÿx2†P0 m/C2/C3 ‡‰n…n‡1†ÿm…m‡1†ŠPmPnˆ0: The first two terms in the last equation can be written as d dx…1ÿx2†…PmP0 nÿPnP0 m†/C2/C3 : Combining this with the last equation we have d dx…1ÿx2†…PmP0 nÿPnP0 m†/C2/C3 ‡‰n…n‡1†ÿm…m‡1†ŠPmPnˆ0: Integrating the above equation between ÿ1 and 1 we obtain …1ÿx2†…PmP0 nÿPnP0 m†j1 ÿ1‡‰n…n‡1†ÿm…m‡1†ŠZ1 ÿ1Pm…x†Pn…x†dxˆ0: The integrated term is zero because (1 ÿx2†ˆ0a txˆ1, and Pm…x†andPn…x† are finite. The bracket in front of the integral is not zero since m6ˆn. Therefore the integral must be zero and we have Z1 ÿ1Pm…x†Pn…x†dxˆ0; m6ˆn: (ii)mˆn: We now use the recurrence relation (7.16a), namely nPn…x†ˆxP0 n…x†ÿP0 nÿ1…x†: Multiplying this recurrence relation by Pn…x†and integrating between ÿ1 and 1, we obtain nZ1 ÿ1Pn…x†‰Š2dxˆZ1 ÿ1xPn…x†P0 n…x†dxÿZ1 ÿ1Pn…x†P0 nÿ1…x†dx: …7:24† The second integral on the right hand side is zero. (Why/C63) To evaluate the first integral on the right hand side, we integrate by parts Z1 ÿ1xPn…x†P0 n…x†dxˆx 2Pn…x†‰Š2j1 ÿ1ÿ1 2Z1 ÿ1Pn…x†‰Š2dxˆ1ÿ12Z 1 ÿ1Pn…x†‰Š2dx: 305LEGENDRE’S EQUATION Substituting these into Eq. (7.24) we obtain nZ1 ÿ1Pn…x†‰Š2dxˆ1ÿ1 2Z1 ÿ1Pn…x†‰Š2dx; which can be simplified to Z1 ÿ1Pn…x†‰Š2dxˆ2 2n‡1: Alternatively, we can use generating function 1 1ÿ2xz‡z2p ˆX1 nˆ0Pn…x†zn: We have on squaring both sides of this: 1 1ÿ2xz‡z2ˆX1 mˆ0X1 nˆ0Pm…x†Pn…x†zm‡n: Then by integrating from ÿ1 to 1 we have Z1 ÿ1dx 1ÿ2xz‡z2ˆX1 mˆ0X1 nˆ0Z1 ÿ1Pm…x†Pn…x†dx/C26/C27 zm‡n: Now Z1 ÿ1dx 1ÿ2xz‡z2ˆÿ1 2zZ1 ÿ1d…1ÿ2xz‡z2† 1ÿ2xz‡z2ˆÿ1 2zln…1ÿ2xz‡z2†j1 ÿ1 and Z1 ÿ1Pm…x†Pn…x†dxˆ0; m6ˆn: Thus, we have ÿ1 2zln…1ÿ2xz‡z2†j1ÿ1ˆX1 nˆ0Z1 ÿ1P2 n…x†dx/C26/C27 z2n or 1 zln1‡z 1ÿz ˆX1 nˆ0Z1 ÿ1P2 n…x†dx/C26/C27 z2n; that is, X1 nˆ02z2n 2n‡1ˆX1 nˆ0Z1 ÿ1P2n…x†dx/C26/C27 z2n: Equating coecients of z2nwe have as requiredR1 ÿ1P2n…x†dxˆ2=…2n‡1†. 306SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Since the Legendre polynomials form a complete orthogonal set on ( ÿ1, 1), we can expand functions in Legendre series just as we expanded functions in Fourier series: f…x†ˆX1 iˆ0ciPi…x†: The coecients cican be found by a method parallel to the one we used in finding the formulas for the coecients in a Fourier series. We shall not pursue this line further. There is a second solution of Legendre’s equation. However, this solution is usually only required in practical applications in which jxj/C621 and we shall only briefly discuss it for such values of x. Now solutions of Legendre’s equation relative to the regular singular point at infinity can be investigated by writingx 2ˆt. With this substitution, dy dxˆdy dtdt dxˆ2t1=2dy dtandd2y dx2ˆd dxdy dx ˆ2dy dx‡4td2y dt2; and Legendre’s equation becomes, after some simplifications, t…1ÿt†d2y dt2‡1 2ÿ32t dy dt‡/C23…/C23‡1† 4yˆ0: This is the hypergeometric equation with ˆÿ/C23=2;/C12ˆ…1‡/C23†=2, and /C13ˆ1 2: x…1ÿx†d2y dx2‡‰/C13ÿ… ‡/C12‡1†xŠdy dxÿ /C12yˆ0; we shall not seek its solutions. The second solution of Legendre’s equation is commonly denoted by /C81/C23…x†and is called the Legendre function of the second kind of order /C23. Thus the general solution of Legendre’s equation (7.1) can be written yˆAP/C23…x†‡B/C81 /C23…x†; AandBbeing arbitrary constants. P/C23…x†is called the Legendre function of the first kind of order /C23and it reduces to the Legendre polynomial Pn…x†when /C23is an integer n. /C84he associated Legendre functions These are the functions of integral order which are solutions of the associatedLegendre equation …1ÿx 2†y00ÿ2xy0‡n…n‡1†ÿm2 1ÿx2() yˆ0 …7:25† with m2n2. 307THE ASSOCIATED LEGENDRE FUNCTIONS We could solve Eq. (7.25) by series; but it is more useful to know how the solutions are related to Legendre polynomials, so we shall proceed in the follow- ing way. We write yˆ…1ÿx2†m=2u…x† and substitute into Eq. (7.25) whence we get, after a little simplification, …1ÿx2†u00ÿ2…m‡1†xu0‡‰n…n‡1†ÿm…m‡1†Šuˆ0: …7:26† Formˆ0, this is a Legendre equation with solution Pn…x†. Now we di/C128erentiate Eq. (7.26) and get …1ÿx2†…u0†00ÿ2‰…m‡1†‡1Šx…u0†0‡‰n…n‡1†ÿ…m‡1†…m‡2†Šu0ˆ0: …7:27† Note that Eq. (7.27) is just Eq. (7.26) with u0in place of u, and ( m‡1) in place of m. Thus, if Pn…x†is a solution of Eq. (7.26) with mˆ0,P0 n…x†is a solution of Eq. (7.26) with mˆ1,P00 n…x†is a solution with mˆ2, and in general for integral m;0mn;…dm=dxm†Pn…x†is a solution of Eq. (7.26). Then yˆ…1ÿx2†m=2dm dxmPn…x†… 7:28† is a solution of the associated Legendre equation (7.25). The functions in Eq.(7.28) are called associated Legendre functions and are denoted by P m n…x†ˆ… 1ÿx2†m=2dm dxmPn…x†: …7:29† Some authors include a factor ( ÿ1†min the definition of Pm n…x†: A negative value of min Eq. (7.25) does not change m2, so a solution of Eq. (7.25) for positive mis also a solution for the corresponding negative m. Thus many references define Pm n…x†forÿnmnas equal to Pjmj n…x†. When we write xˆcos, Eq. (7.25) becomes 1 sind dsindy d ‡n…n‡1†ÿm2 sin2() yˆ0 …7:30† and Eq. (7.29) becomes Pm n…cos†ˆsinmdm d…cos†mPn…cos† fg : In particular Dÿ1meansZx 1Pn…x†dx: 308SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Orthogonality of associated /C76egendre functions As in the case of Legendre polynomials, the associated Legendre functions Pm n…x† are orthogonal for ÿ1x1 and in particular Z1 ÿ1Ps m…x†Ps n…x†dxˆ…n‡s†/C33 …nÿs†/C33mn: …7:31† To prove this, let us write for simplicity MˆPm s…x†;and /C78ˆPsn…x† and from Eq. (7.25), the associated Legendre equation, we have d dx…1ÿx2†dM dx/C26/C27 ‡m…m‡1†ÿs2 1ÿx2() Mˆ0 …7:32† and d dx…1ÿx2†d/C78 dx/C26/C27 ‡n…n‡1†ÿs2 1ÿx2() /C78ˆ0: …7:33† Multiplying Eq. (7.32) by /C78, Eq. (7.33) by Mand subtracting, we get Md dx…1ÿx2†d/C78 dx/C26/C27 ÿ/C78d dx…1ÿx2†dM dx/C26/C27 ˆfm…m‡1†ÿn…n‡1†gM/C78: Integration between ÿ1 and 1 gives …mÿn†…m‡nÿ1†Z1 ÿ1M/C78dx ˆZ1 ÿ1 Md dx…1ÿx2†d/C78 dx/C26/C27 ÿ/C78d dx…1ÿx2†dM dx/C26/C27  dx:…7:34† Integration by parts gives Z1 ÿ1Md dxf…1ÿx2†/C780gdxˆ‰M/C780…1ÿx2†Š1 ÿ1ÿZ1 ÿ1…1ÿx2†M0/C780dx ˆÿZ1 ÿ1…1ÿx2†M0/C780dx: 309THE ASSOCIATED LEGENDRE FUNCTIONS Then integrating by parts once more, we obtain Z1 ÿ1Md dxf…1ÿx2†/C780gdxˆÿZ1 ÿ1…1ÿx2†M0/C780dx ˆÿ ‰M/C78…1ÿx2†Š1 ÿ1‡Z1 ÿ1/C78d dx…1ÿx2†M0/C8/C9 dx ˆZ1 ÿ1/C78d dxf…1ÿx2†M0gdx: Substituting this in Eq. (7.34) we get …mÿn†…m‡nÿ1†Z1 ÿ1M/C78dx ˆ0: Ifmn, we have Z1 ÿ1M/C78dx ˆZ1 ÿ1Ps m…x†Psm…x†dxˆ0 …m6ˆn†: Ifmˆn, let us write Ps n…x†ˆ… 1ÿx2†s=2ds dxsPn…x†ˆ…1ÿx2†s=2 2nn/C33ds‡n dxs‡nf…x2ÿ1†ng: Hence Z1 ÿ1Ps n…x†Psn…x†dxˆ1 22n…n/C33†2Z1 ÿ1…1ÿx2†sDn‡sf…x2ÿ1†ngDn‡s f …x2ÿ1†ngdx; …Dkˆdk=dxk†: Integration by parts gives 1 22n…n/C33†2‰…1ÿx2†sDn‡sf…x2ÿ1†ngDn‡sÿ1f…x2ÿ1†ngŠ1 ÿ1 ÿ1 22n…n/C33†2Z1 ÿ1D…1ÿx2†sDn‡s…x2ÿ1†n/C8/C9 /C2/C3 Dn‡sÿ1…x2ÿ1†n/C8/C9 dx: The first term vanishes at both limits and we have Z1 ÿ1fPs n…x†g2dxˆÿ1 22n…n/C33†2Z1 ÿ1D‰…1ÿx2†sDn‡sf…x2ÿ1†ngŠDn‡sÿ1f…x2ÿ1†ngdx: …7:35† We can continue to integrate Eq. (7.35) by parts and the first term continues to vanish since D/C112‰…1ÿx2†sDn‡sf…x2ÿ1†ngŠcontains the factor …1ÿx2†when /C112<s 310SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS andDn‡sÿ/C112f…x2ÿ1†ngcontains it when /C112s. After integrating …n‡s†times we find Z1 ÿ1fPs n…x†g2dxˆ…ÿ1†n‡s 22n…n/C33†2Z1 ÿ1Dn‡s‰…1ÿx2†sDn‡sf…x2ÿ1†ngŠ…x2ÿ1†ndx:…7:36† But Dn‡sf…x2ÿ1†ngis a polynomial of degree ( nÿs) so that (1ÿx2†sDn‡sf…x2ÿ1†ngis of degree nÿ2‡2sˆn‡s. Hence the first factor in the integrand is a polynomial of degree zero. We can find this constant by examining the following: Dn‡s…x2n†ˆ2n…2nÿ1†…2nÿ2†… nÿ‡1†xnÿs: Hence the highest power in …1ÿx2†sDn‡sf…x2ÿ1†ngis the term …ÿ1†s2n…2nÿ1†… nÿs‡1†xn‡s; so that Dn‡s‰…1ÿx2†sDn‡sf…x2ÿ1†ngŠ ˆ …ÿ 1†s…2n†/C33…n‡s†/C33 …nÿs†/C33: Now Eq. (7.36) gives, by writing xˆcos, Z1 ÿ1Ps nf…x†g2dxˆ…ÿ1†n 22n…n/C33†2Z1 ÿ1…2n†/C33…n‡s†/C33 …nÿs†/C33…x2ÿ1†ndx ˆ2 2n‡1…n‡s†/C33 …nÿs†/C33…7:37† /C72ermite/C39s equation Hermite’s equation is y00ÿ2xy0‡2/C23yˆ0; …7:38† where y0ˆdy=dx. The reader will see this equation in quantum mechanics (when solving the Schro /C200dinger equation for a linear harmonic potential function). The origin xˆ0 is an ordinary point and we may write the solution in the form yˆa0‡a1x‡a2x2‡ˆX1 jˆ0ajxj: …7:39† Di/C128erentiating the series term by term, we have y0ˆX1 jˆ0jajxjÿ1; y00ˆX1 jˆ0…j‡1†…j‡2†aj‡2xj: 311HERMITE’S EQUATION Substituting these into Eq. (7.38) we obtain X1 jˆ0…j‡1†…j‡2†aj‡2‡2…/C23ÿj†aj/C2/C3 xjˆ0: For a power series to vanish the coecient of each power of xmust be zero; this gives …j‡1†…j‡2†aj‡2‡2…/C23ÿj†ajˆ0; from which we obtain the recurrence relations aj‡2ˆ2…jÿ/C23† …j‡1†…j‡2†aj: …7:40† We obtain polynomial solutions of Eq. (7.38) when /C23ˆn, a positive integer. Then Eq. (7.40) gives an‡2ˆan‡4ˆˆ 0: For even n, Eq. (7.40) gives a2ˆ… ÿ 1†2n 2/C33a0;a4ˆ… ÿ 1†222…nÿ2†n 4/C33a0;a6ˆ… ÿ 1†323…nÿ4†…nÿ2†n 6/C33a0 and generally anˆ… ÿ 1†n=22n=2n…nÿ2†42 n/C33a0: This solution is called a Hermite polynomial of degree nand is written Hn…x†.I f we choose a0ˆ…ÿ1†n=22n=2n/C33 n…nÿ2†42ˆ…ÿ1†n=2n/C33 …n=2†/C33 we can write Hn…x†ˆ… 2x†nÿn…nÿ1† 1/C33…2x†nÿ2‡n…nÿ1†…nÿ2†…nÿ3† 2/C33…2x†nÿ4‡ :…7:41† When nis odd the polynomial solution of Eq. (7.38) can still be written as Eq. (7.41) if we write a1ˆ…ÿ1†…nÿ1†=22n/C33 …n=2ÿ1=2†/C33: In particular, H0…x†ˆ1;H1…x†ˆ2x;H3…x†ˆ4x2ÿ2;H3…x†ˆ8x2ÿ12x; H4…x†ˆ16x4ÿ48x2‡12;H5…x†ˆ32x5ÿ160x3‡120x;...: 312SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Rodrigues’ formula for /C72ermite polynomials /C72n…x† The Hermite polynomials are also given by the formula Hn…x†ˆ… ÿ 1†nex2dn dxn…eÿx2†: …7:42† To prove this formula, let us write /C113ˆeÿx2. Then D/C113‡2x/C113ˆ0;Dˆd dx: Di/C128erentiate this ( n‡1) times by the Leibnitz’ rule giving Dn‡2/C113‡2xDn‡1/C113‡2…n‡1†Dn/C113ˆ0: Writing yˆ… ÿ 1†nDn/C113gives D2y‡2xDy ‡2…n‡1†yˆ0 …7:43† substitute uˆex2ythen Duˆex2f2xy‡Dyg and D2uˆex2fD2y‡4xDy ‡4x2y‡2yg: Hence by Eq. (7.43) we get D2uÿ2xDu ‡2nuˆ0; which indicates that uˆ… ÿ 1†nex2Dn…eÿx2† is a polynomial solution of Hermite’s equation (7.38). Recurrence relations for /C72ermite polynomials Rodrigues’ formula gives on di/C128erentiation H0 n…x†ˆ… ÿ 1†n2xex2Dn…eÿx2†‡… ÿ 1†nex2Dn‡1…eÿx2†: that is, H0 n…x†ˆ2xHn…x†ÿHn‡1…x†: …7:44† Eq. (7.44) gives on di/C128erentiation H00 n…x†ˆ2Hn…x†‡2xH0 n…x†ÿH0 n‡1…x†: 313HERMITE’S EQUATION Now Hn…x†satisfies Hermite’s equation H00 n…x†ÿ2xH0 n…x†‡2nHn…x†ˆ0: Eliminating H00 n…x†from the last two equations, we obtain 2xH0 n…x†ÿ2nHn…x†ˆ2Hn…x†‡2xH0 n…x†ÿH0 n‡1…x† which reduces to H0 n‡1…x†ˆ2…n‡1†Hn…x†: …7:45† Replacing nbyn‡1 in Eq. (7.44), we have H0 n‡1…x†ˆ2xHn‡1…x†ÿHn‡2…x†: Combining this with Eq. (7.45) we obtain Hn‡2…x†ˆ2xHn‡1…x†ÿ2…n‡1†Hn…x†: …7:46† This will quickly give the higher polynomials. Generating function for the /C72n…x† By using Rodrigues’ formula we can also find a generating formula for the Hn…x†. This is …x;t†ˆe2txÿt2ˆefx2ÿ…tÿx†2gˆX1 nˆ0Hn…x† n/C33tn: …7:47† Di/C128erentiating Eq. (7.47) ntimes with respect to twe get ex2/C64n /C64tneÿ…tÿx†2ˆex2…ÿ1†n/C64n /C64xneÿ…tÿx†2ˆX1 kˆ0Hn‡k…x†tk k/C33: Puttˆ0 in the last equation and we obtain Rodrigues’ formula Hn…x†ˆ… ÿ 1†nex2dn dxn…eÿx2†: /C84he orthogonal /C72ermite functions These are defined by /C70n…x†ˆeÿx2=2Hn…x†; …7:48† from which we have D/C70n…x†ˆÿ x/C70n…x†‡eÿx2=2H0 n…x†; D2/C70n…x†ˆeÿx2=2H00 n…x†ÿ2xeÿx2=2H0 n…x†‡x2eÿx2=2Hn…x†ÿ/C70n…x† ˆeÿx2=2‰H00 n…x†ÿ2xH0 n…x†Š ‡x2/C70n…x†ÿ/C70n…x†; 314SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS butH00 n…x†ÿ2xH0 n…x†ˆÿ 2nHn…x†, so we can rewrite the last equation as D2/C70n…x†ˆeÿx2=2‰ÿ2nH0 n…x†Š ‡x2/C70n…x†ÿ/C70n…x† ˆÿ2n/C70n…x†‡x2/C70n…x†ÿ/C70n…x†; which gives D2/C70n…x†ÿx2/C70n…x†‡…2n‡1†/C70n…x†ˆ0: …7:49† We can now show that the set f/C70n…x†gis orthogonal in the infinite range ÿ1 <x<1. Multiplying Eq. (7.49) by /C70m…x†we have /C70m…x†D2/C70n…x†ÿx2/C70n…x†/C70m…x†‡…2n‡1†/C70n…x†/C70m…x†ˆ0: Interchanging mandngives /C70n…x†D2/C70m…x†ÿx2/C70m…x†/C70n…x†‡…2m‡1†/C70m…x†/C70n…x†ˆ0: Subtracting the last two equations from the previous one and then integrating from ÿ1 to‡1,w eh a v e In;mˆZ1 ÿ1/C70n…x†/C70m…x†dxˆ1 2…nÿm†Z1 ÿ1…/C7000 n/C70mÿ/C7000 m/C70n†dx: The integration by parts gives 2…nÿm†In;mˆ/C700 n/C70mÿ/C700 m/C70n/C2/C31 ÿ1ÿZ1 ÿ1…/C700 n/C700 mÿ/C700 m/C700 n†dx: Since the right hand side vanishes at both limits and if m6ˆm, we have In;mˆZ1 ÿ1/C70n…x†/C70m…x†dxˆ0: …7:50† When nˆmwe can proceed as follows In;nˆZ1 ÿ1eÿx2Hn…x†Hn…x†dxˆZ1 ÿ1ex2Dn…eÿx2†Dm…eÿx2†dx: Integration by parts, that is,R ud/C118ˆu/C118ÿR /C118du with uˆeÿx2Dn…eÿx2† and/C118ˆDnÿ1…eÿx2†, gives In;nˆÿZ1 ÿ1‰2xex2Dn…eÿx2†‡ex2Dn‡1…eÿx2†ŠDnÿ1…eÿx2†dx: By using Eq. (7.43) which is true for yˆ… ÿ 1†nDn/C113ˆ… ÿ 1†nDn…eÿx2†we obtain In;nˆZ1 ÿ12nex2Dnÿ1…eÿx2†Dnÿ1…eÿx2†dxˆ2nInÿ1;nÿ1: 315HERMITE’S EQUATION Since I0;0ˆZ1 ÿ1eÿx2dxˆÿ…1=2†ˆp; we find that In;nˆZ1 ÿ1eÿx2Hn…x†Hn…x†dxˆ2nn/C33p: …7:51† We can also use the generating function for the Hermite polynomials: e2txÿt2ˆX1 nˆ0Hn…x†tn n/C33;e2sxÿs2ˆX1 mˆ0Hm…x†sm m/C33: Multiplying these, we have e2txÿt2‡2sxÿs2ˆX1 mˆ0X1 nˆ0Hm…x†Hn…x†smtn m/C33n/C33: Multiplying by eÿx2and integrating from ÿ1 to1gives Z1 ÿ1eÿ‰…x‡s‡t†2ÿ2stŠdxˆX1 mˆ0X1 nˆ0smtn m/C33n/C33Z1 ÿ1eÿx2Hm…x†Hn…x†dx: Now the left hand side is equal to e2stZ1 ÿ1eÿ…x‡s‡t†2dxˆe2stZ1 ÿ1eÿu2duˆe2stpˆpX 1 mˆ02msmtm m/C33: By equating coecients the required result follows. It follows that the functions …1=2nn/C33np†1=2eÿx2Hn…x†form an orthonormal set. We shall assume it is complete. Laguerre/C39s equation Laguerre’s equation is xD2y‡…1ÿx†Dy‡/C23yˆ0: …7:52† This equation and its solutions (Laguerre functions) are of interest in quantum mechanics (e.g., the hydrogen problem). The origin xˆ0 is a regular singular point and so we write y…x†ˆX1 kˆ0akxk‡/C26: …7:53† By substitution, Eq. (7.52) becomes X1 kˆ0‰…k‡/C26†2akxk‡/C26ÿ1‡…/C23ÿk‡/C26†akxkŠˆ0 …7:54† 316SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS from which we find that the indicial equation is /C262ˆ0. And then (7.54) reduces to X1 kˆ0‰k2akxkÿ1‡…/C23ÿk†akxkŠˆ0: Changing kÿ1t ok0in the first term, then renaming k0ˆk, we obtain X1 kˆ0f…k‡1†2ak‡1‡…/C23ÿk†akgxkˆ0; whence the recurrence relations are ak‡1ˆkÿ/C23 …k‡1†2ak: …7:55† When /C23is a positive integer n, the recurrence relations give ak‡1ˆak‡2ˆˆ 0, and a1ˆÿn 12a0; a2ˆÿ…nÿ1† 22a1ˆ…ÿ1†2…nÿ1†n …12†2a0; a3ˆÿ…nÿ2† 32a2ˆ…ÿ1†3…nÿ2†…nÿ1†n …123†2a0;etc: In general akˆ… ÿ 1†k…nÿk‡1†…nÿk‡2†… nÿ1†n …k/C33†2a0: …7:56† We usually choose a0ˆ… ÿ 1†n/C33, then the polynomial solution of Eq. (7.52) is given by Ln…x†ˆ… ÿ 1†nxnÿn2 1/C33xnÿ1‡n2…nÿ1†2 2/C33xnÿ2ÿ‡‡… ÿ 1†nn/C33() :…7:57† This is called the Laguerre polynomial of degree n. We list the first four Laguerre polynomials below: L0…x†ˆ1;L1…x†ˆ1ÿx;L2…x†ˆ2ÿ4x‡x2;L3…x†ˆ6ÿ18x‡9x2ÿx3: /C84he generating function for the /C76aguerre polynomials Ln…x† This is given by …x;z†ˆeÿxz=…1ÿz† 1ÿzˆX1 nˆ0Ln…x† n/C33zn: …7:58† 317LAGUERRE’S EQUATION By writing the series for the exponential and collecting powers of z, you can verify the first few terms of the series. And it is also straightforward to show that x/C642 /C64x2‡…1ÿx†/C64 /C64x‡z/C64 /C64zˆ0: Substituting the right hand side of Eq. (7.58), that is, …x;z†ˆP1 nˆ0‰Ln…x†=n/C33Šzn, into the last equation we see that the functions Ln…x†satisfy Laguerre’s equation. Thus we identify …x;z†as the generating function for the Laguerre polynomials. Now multiplying Eq. (7.58) by zÿnÿ1and integrating around the origin, we obtain Ln…x†ˆn/C33 2iIeÿxz=…1ÿz† …1ÿz†zn‡1dz; …7:59† which is an integral representation of Ln…x†. By di/C128erentiating the generating function in Eq. (7.58) with respect to xandz, we obtain the recurrence relations Ln‡1…x†ˆ… 2n‡1ÿx†Ln…x†ÿn2Lnÿ1…x†; nLnÿ1…x†ˆnL0 nÿ1…x†ÿL0 n…x†:) …7:60† Rodrigues’ formula for the /C76aguerre polynomials Ln…x† The Laguerre polynomials are also given by Rodrigues’ formula Ln…x†ˆexdn dxn…xneÿx†: …7:61† To prove this formula, let us go back to the integral representation of Ln…x†, Eq. (7.59). With the transformation xz 1ÿzˆsÿxorzˆsÿx s; Eq. (7.59) becomes Ln…x†ˆn/C33ex 2iIsneÿn …sÿx†n‡1ds; the new contour enclosing the point sˆxin the splane. By Cauchy’s integral formula (for derivatives) this reduces to Ln…x†ˆexdn dxn…xneÿx†; which is Rodrigues’ formula. Alternatively, we can di/C128erentiate Eq. (7.58) ntimes with respect to zand afterwards put z/C610, and thus obtain exlim z!0/C64n /C64zn…1ÿz†ÿ1expÿx 1ÿz /C104/C105 ˆLn…x†: 318SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS But lim z!0/C64n /C64zn…1ÿz†ÿ1expÿx 1ÿz /C104/C105 ˆdn dxnxneÿx…† ; hence Ln…x†ˆexdn dxn…xneÿx†: /C84he orthogonal /C76aguerre functions The Laguerre polynomials, Ln…x†, do not by themselves form an orthogonal set. But the functions eÿx=2Ln…x†are orthogonal in the interval (0, 1). For any two Laguerre polynomials Lm…x†andLn…x†we have, from Laguerre’s equation, xL00 m‡…1ÿx†L0 m‡mLmˆ0; xL00 n‡…1ÿx†L0 n‡mLnˆ0: Multiplying these equations by Ln…x†andLm…x†respectively and subtracting, we find x‰LnL00 mÿLmL00 nЇ…1ÿx†‰LnL0 mÿLmL0 nŠˆ…nÿm†LmLn or d dx‰LnL0 mÿLmL0 nЇ1ÿx x‰LnL0 mÿLmL0 nŠˆ…nÿm†LmLn x: Then multiplying by the integrating factor expZ ‰…1ÿx†=xŠdxˆexp…lnxÿx†ˆxeÿx; we have d dxfxeÿx‰LnL0 mÿLmL0 nŠg ˆ … nÿm†eÿxLmLn: Integrating from 0 to 1gives …nÿm†Z1 0eÿxLm…x†Ln…x†dxˆxeÿx‰LnL0 mÿLmL0 nŠj1 0ˆ0: Thus if m6ˆn Z1 0eÿxLm…x†Ln…x†dxˆ0…m6ˆn†; …7:62† which proves the required result. Alternatively, we can use Rodrigues’ formula (7.61). If mis a positive integer, Z1 0eÿxxmLm…x†dxˆZ1 0xmdn dxn…xneÿx†dxˆ… ÿ 1†mm/C33Z1 0dnÿm dxnÿm…xneÿx†dx; …7:63† 319LAGUERRE’S EQUATION the last step resulting from integrating by parts mtimes. The integral on the right hand side is zero when n/C62mand, since Ln…x†is a polynomial of degree minx,i t follows that Z1 0eÿxLm…x†Ln…x†dxˆ0…m6ˆn†; which is Eq. (7.62). The reader can also apply Eq. (7.63) to show that Z1 0eÿxLn…x† fg2dxˆ…n/C33†2: …7:64† Hence the functions feÿx=2Ln…x†=n/C33gform an orthonormal system. /C84he associated Laguerre pol/C121nomials Lm n…x† Di/C128erentiating Laguerre’s equation (7.52) mtimes by the Leibnitz theorem we obtain xDm‡2y‡…m‡1ÿx†Dm‡1y‡…nÿm†Dmyˆ0…/C23ˆn† and writing zˆDmywe obtain xD2z‡…m‡1ÿx†Dz‡…nÿm†zˆ0: …7:65† This is Laguerre’s associated equation and it clearly possesses a polynomial solu- tion zˆDmLn…x†Lm n…x†…mn†; …7:66† called the associated Laguerre polynomial of degree ( nÿm). Using Rodrigues’ formula for Laguerre polynomial Ln…x†, Eq. (7.61), we obtain Lmn…x†ˆdm dxmLn…x†ˆdm dxmexdn dxn…xneÿx†/C26/C27 : …7:67† This result is very useful in establishing further properties of the associated Laguerre polynomials. The first few polynomials are listed below: L0 0…x†ˆ1;L01…x†ˆ1ÿx;L11…x†ˆÿ 1; L02…x†ˆ2ÿ4x‡x2;L12…x†ˆÿ 4‡2x;L22…x†ˆ2: Generating function for the associated /C76aguerre polynomials The Laguerre polynomial Ln…x†can be generated by the function 1 1ÿtexpÿxt 1ÿt ˆX1 nˆ0Ln…x†tn n/C33: 320SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Di/C128erentiating this ktimes with respect to x, it is seen at once that …ÿ1†k…1ÿt†ÿ1t 1ÿt k expÿxt 1ÿt ˆX1 ˆkLk …x† /C33t: …7:68† /C65ssociated /C76aguerre function of integral order A function of great importance in quantum mechanics is the associated Laguerre function that is defined as /C71m n…x†ˆeÿx=2x…mÿ1†=2Lmn…x†…mn†: …7:69† It is significant largely because j/C71m n…x†j ! 0a sx!1 . It satisfies the di/C128erential equation x2D2u‡2xDu‡ nÿmÿ1 2 xÿx2 4ÿm2ÿ1 4"# uˆ0: …7:70† If we substitute uˆeÿx=2x…mÿ1†=2zin this equation, it reduces to Laguerre’s asso- ciated equation (7.65). Thus uˆ/C71mnsatisfies Eq. (7.70). You will meet this equa- tion in quantum mechanics in the study of the hydrogen atom. Certain integrals involving /C71mnare often used in quantum mechanics and they are of the form In;mˆZ1 0eÿxxkÿ1Lkn…x†Lkm…x†x/C112dx; where pis also an integer. We will not consider these here and instead refer the interested reader to the following book: The Mathematics of Physics and /C67hemistry , by Henry Margenau and George M. Murphy; D. Van Nostrand Co. Inc., New York, 1956. /C66essel/C39s equation The di/C128erential equation x2y00‡xy0‡…x2ÿ 2†yˆ0 …7:71† in which is a real and positive constant, is known as Bessel’s equation and its solutions are called Bessel functions. These functions were used by Bessel (Friedrich Wilhelm Bessel, 1784–1864, German mathematician and astronomer) extensively in a problem of dynamical astronomy. The importance of this equation and its solutions (Bessel functions) lies in the fact that they occur fre- quently in the boundary-value problems of mathematical physics and engineering 321BESSEL’S EQUATION involving cylindrical symmetry (so Bessel functions are sometimes called cylind- rical functions), and many others. There are whole books on Bessel functions. The origin is a regular singular point, and all other values of x are ordinary points. At the origin we seek a series solution of the form y…x†ˆX1 mˆ0amxm‡/C26…a06ˆ0†: …7:72† Substituting this and its derivatives into Bessel’s equation (7.71), we have X1 mˆ0…m‡/C26†…m‡/C26ÿ1†amxm‡/C26‡X1 mˆ0…m‡/C26†amxm‡/C26 ‡X1 mˆ0amxm‡/C26‡2ÿ 2X1 mˆ0amxm‡/C26ˆ0: This will be an identity if and only if the coecient of every power of xis zero. By equating the sum of the coecients of xk‡/C26to zero we find /C26…/C26ÿ1†a0‡/C26a0ÿ 2a0ˆ0 …kˆ0†; …7:73a† …/C26ÿ1†/C26a1‡…/C26‡1†a1ÿ 2a1ˆ0 …kˆ1†; …7:73b† …k‡/C26†…k‡/C26ÿ1†ak‡…k‡/C26†ak‡akÿ2ÿ 2akˆ0 …kˆ2;3;...†:…7:73c† From Eq. (7.73a) we obtain the indicial equation /C26…/C26ÿ1†‡/C26ÿ 2ˆ…/C26‡ †…/C26ÿ †ˆ0: The roots are /C26ˆ . We first determine a solution corresponding to the positive root. For /C26ˆ‡ , Eq. (7.73b) yields a1ˆ0, and Eq. (7.73c) takes the form …k‡2 †kak‡akÿ2ˆ0; or akˆÿ1 k…k‡2 †akÿ2; …7:74† which is a recurrence formula: since a1ˆ0 and 0, it follows that a3ˆ0;a5ˆ0;...;successively. If we set kˆ2min Eq. (7.74), the recurrence formula becomes a2mˆÿ1 22m… ‡m†a2mÿ2; mˆ1;2;... …7:75† and we can determine the coecients a2;a4, successively. We can rewrite a2min terms of a0: a2mˆ…ÿ1†m 22mm/C33… ‡m†… ‡2†… ‡1†a0: 322SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Now a2mis the coecient of x ‡2min the series (7.72) for y. Hence it would be convenient if a2mcontained the factor 2 ‡2min its denominator instead of just 22m. To achieve this, we write a2mˆ…ÿ1†m 2 ‡2mm/C33… ‡m†… ‡2†… ‡1†…2 a0†: Furthermore, the factors … ‡m†… ‡2†… ‡1† suggest a factorial. In fact, if were an integer, a factorial could be created by multiplying numerator by /C33. However, since is not necessarily an integer, we must use not /C33but its generalization ÿ… ‡1†for this purpose. Then, except for the values ˆÿ1;ÿ2;ÿ3;... for which ÿ… ‡1†is not defined, we can write a2mˆ…ÿ1†m 2 ‡2mm/C33… ‡m†… ‡2†… ‡1†ÿ… ‡1†‰2 ÿ… ‡1†a0Š: Since the gamma function satisfies the recurrence relation zÿ…z†ˆÿ…z‡1†, the expression for a2mbecomes finally a2mˆ…ÿ1†m 2 ‡2mm/C33ÿ… ‡m‡1†‰2 ÿ… ‡1†a0Š: Since a0is arbitrary, and since we are looking only for particular solutions, we choose a0ˆ1 2 ÿ… ‡1†; so that a2mˆ…ÿ1†m 2 ‡2mm/C33ÿ… ‡m‡1†; a2m‡1ˆ0 and the series for yis, from Eq. (7.72), y…x†ˆx 1 2 ÿ… ‡1†ÿx2 2 ‡2ÿ… ‡2†‡x4 2 ‡42/C33ÿ… ‡3†ÿ‡"# ˆX1 mˆ0…ÿ1†m 2 ‡2mm/C33ÿ… ‡m‡1†x ‡2m: …7:76† The function defined by this infinite series is known as the Bessel function of the first kind of order and is denoted by the symbol /C74 …x†. Since Bessel’s equation 323BESSEL’S EQUATION of order has no finite singular points except the origin, the ratio test will show that the series for /C74 …x†converges for all values of xif 0. When ˆn, an integer, solution (7.76) becomes, for n0 /C74n…x†ˆxnX1 mˆ0…ÿ1†mx2m 22m‡nm/C33…n‡m†/C33: …7:76a† The graphs of /C740…x†;/C741…x†, and /C742…x†are shown in Fig. 7.3. Their resemblance to the graphs of cos xand sin xis interesting (Problem 7.16 illustrates this for the first few terms). Fig. 7.3 also illustrates the important fact that for every value of the equation /C74 …x†ˆ0 has infinitely many real roots. With the second root /C26ˆÿ of the indicial equation, the recurrence relation takes the form (from Eq. (7.73c)) akˆÿ1 k…kÿ2 †akÿ2: …7:77† If is not an integer, this leads to an independent second solution that can be written /C74ÿ …x†ˆX1 mˆ0…ÿ1†m m/C33ÿ…ÿ ‡m‡1†…x=2†ÿ ‡2m…7:78† and the complete solution of Bessel’s equation is then y…x†ˆA/C74 …x†‡B/C74ÿ …x†; …7:79† where AandBare arbitrary constants. When is a positive integer n, it can be shown that the formal expression for /C74ÿn…x†is equal to ( ÿ1†n/C74n…x†.S o/C74n…x†and/C74ÿn…x†are linearly dependent and Eq. (7.79) cannot be a general solution. In fact, if is a positive integer, the recurrence 324SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Figure 7.3. Bessel functions of the first kind. relation (7.77) breaks down when 2 ˆkand a second solution has to be found by other methods. There is a diculty also when ˆ0, in which case the two roots of the indicial equation are equal; the second solution must also found by other methods. These will be discussed in next section. The results of Problem 7.16 are a special case of an important general theorem which states that /C74 …x†is expressible in finite terms by means of algebraic and trigonometrical functions of xwhenever is half of an odd integer. Further examples are /C743=2…x†ˆ2 x1=2sinx xÿcosx ; /C74ÿ5=2…x†ˆ2 x1=23 sinx x‡3 x2ÿ1 cosx/C26/C27 : The functions /C74…n‡1=2†…x†and/C74ÿ…n‡1=2†…x†, where nis a positive integer or zero, are called spherical Bessel functions; they have important applications in problems of wave motion in which spherical polar coordinates are appropriate. Bessel functions of the second /C107ind /C89n…x† For integer ˆn;/C74n…x†and/C74ÿn…x†are linearly dependent and do not form a fundamental system. We shall now obtain a second independent solution, startingwith the case nˆ0. In this case Bessel’s equation may be written xy 00‡y0‡xyˆ0; …7:80† the indicial equation (7.73a) now, with ˆ0, has the double root /C26ˆ0. Then we see from Eq. (7.33) that the desired solution must be of the form y2…x†ˆ/C740…x†lnx‡X1 mˆ1Amxm: …7:81† Next we substitute y2and its derivatives y0 2ˆ/C740 0lnx‡/C740 x‡X1 mˆ1mA mxmÿ1; y00 2ˆ/C7400 0lnx‡2/C740 0 xÿ/C740 x2‡X1 mˆ1m…mÿ1†Amxmÿ2 into Eq. (7.80). Then the logarithmic terms disappear because /C740is a solution of Eq. (7.80), the other two terms containing /C740cancel, and we find 2/C740 0‡X1 mˆ1m…mÿ1†Amxmÿ1‡X1 mˆ1mA mxmÿ1‡X1 mˆ1Amxm‡1ˆ0: 325BESSEL’S EQUATION From Eq. (7.76a) we obtain /C740 0as /C740 0…x†ˆX1 mˆ1…ÿ1†m2mx2mÿ1 22m…m/C33†2ˆX1 mˆ1…ÿ1†mx2mÿ1 22mÿ1m/C33…mÿ1†/C33: By inserting this series we have X1 mˆ1…ÿ1†mx2mÿ1 22mÿ2m/C33…mÿ1†/C33‡X1 mˆ1m2Amxmÿ1‡X1 mˆ1Amxm‡1ˆ0: We first show that Amwith odd subscripts are all zero. The coecient of the power x0isA1and so A1ˆ0. By equating the sum of the coecients of the power x2sto zero we obtain …2s‡1†2A2s‡1‡A2sÿ1ˆ0; sˆ1;2;...: Since A1ˆ0, we thus obtain A3ˆ0;A5ˆ0;...;successively. We now equate the sum of the coecients of x2s‡1to zero. For sˆ0 this gives ÿ1‡4A2ˆ0o r A2ˆ1=4: For the other values of swe obtain …ÿ1†s‡1 2s…s‡1†/C33s/C33‡…2s‡2†2A2s‡2‡A2sˆ0: Forsˆ1 this yields 1=8‡16A4‡A2ˆ0o r A4ˆÿ3=128 and in general A2mˆ…ÿ1†mÿ1 2m…m/C33†21‡1 2‡13‡‡1 m ; mˆ1;2;...: …7:82† Using the short notation /C104mˆ1‡1 2‡13‡‡1 m and inserting Eq. (7.82) and A1ˆA3ˆˆ 0 into Eq. (7.81) we obtain the result y2…x†ˆ/C740…x†lnx‡X1 mˆ1…ÿ1†mÿ1/C104m 22m…m/C33†2x2m ˆ/C740…x†lnx‡14x 2ÿ3 128x4‡ÿ : …7:83† Since /C740andy2are linearly independent functions, they form a fundamental system of Eq. (7.80). Of course, another fundamental system is obtained by replacing y2by an independent particular solution of the form a…y2‡b/C740†, where a…6ˆ0†and bare constants. It is customary to choose aˆ2=and 326SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS bˆ/C13ÿln 2, where /C13ˆ0:577 215 664 90 ...is the so-called Euler constant, which is defined as the limit of 1‡1 2‡‡1 sÿlns assapproaches infinity. The standard particular solution thus obtained is known as the Bessel function of the second kind of order zero or Neumann’s function of order zero and is denoted by Y0…x†: Y0…x†ˆ2  /C740…x†lnx 2‡/C13 ‡X1 mˆ1…ÿ1†mÿ1/C104m 22m…m/C33†2x2m : …7:84† If ˆ1;2;...;a second solution can be obtained by similar manipulations, starting from Eq. (7.35). It turns out that in this case also the solution contains a logarithmic term. So the second solution is unbounded near the origin and is useful in applications only for x6ˆ0. Note that the second solution is defined di/C128erently, depending on whether the order is integral or not. To provide uniformity of formalism and numerical tabulation, it is desirable to adopt a form of the second solution that is valid forall values of the order. The common choice for the standard second solution defined for all is given by the formula Y …x†ˆ/C74 …x†cos ÿ/C74ÿ …x† sin ;Yn…x†ˆlim !nY …x†: …7:85† This function is known as the Bessel function of the second kind of order .I ti s also known as Neumann’s function of order and is denoted by /C78 …x†(Carl Neumann 1832–1925, German mathematician and physicist). In G. N. Watson’sA Treatise on the Theory of Bessel /C70unctions (2nd ed. Cambridge University Press, Cambridge, 1944), it was called Weber’s function and the notation Y …x†was used. It can be shown that Yÿn…x†ˆ… ÿ 1†nYn…x†: We plot the first three Yn…x†in Fig. 7.4. A general solution of Bessel’s equation for all values of can now be written: y…x†ˆc1/C74 …x†‡c2Y …x†: In some applications it is convenient to use solutions of Bessel’s equation that are complex for all values of x, so the following solutions were introduced H…1† …x†ˆ/C74 …x†‡iY …x†; H…2† …x†ˆ/C74 …x†ÿiY …x†:9 = ;…7:86† 327BESSEL’S EQUATION These linearly independent functions are known as Bessel functions of the third kind of order or first and second Hankel functions of order (Hermann Hankel, 1839–1873, German mathematician). To illustrate how Bessel functions enter into the analysis of physical problems, we consider one example in classical physics: small oscillations of a hanging chain, which was first considered as early as 1732 by Daniel Bernoulli. /C72anging /C175exible chain Fig. 7.5 shows a uniform heavy flexible chain of length lhanging vertically under its own weight. The x-axis is the position of stable equilibrium of the chain and its lowest end is at xˆ0. We consider the problem of small oscillations in the vertical xyplane caused by small displacements from the stable equilibrium position. This is essentially the problem of the vibrating string which we discussed in Chapter 4, with two important di/C128erences: here, instead of being constant, the tension Tat a given point of the chain is equal to the weight of the chain below that point, andnow one end of the chain is free, whereas before both ends were fixed. The analysis of Chapter 4 generally holds. To derive an equation for y, consider an element dx, then Newton’s second law gives T/C64y /C64x 2ÿT/C64y /C64x 1ˆ/C26dx/C642y /C64t2 or /C26dx/C642y /C64t2ˆ/C64 /C64xT/C64y /C64x dx; 328SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Figure 7.4. Bessel functions of the second kind. from which we obtain /C26/C642y /C64t2ˆ/C64 /C64xT/C64y /C64x : Now Tˆ/C26/C103x. Substituting this into the above equation for y, we obtain /C642y /C64t2ˆ/C103/C64y /C64x‡/C103x/C642y /C64x2; where yis a function of two variables xandt. The first step in the solution is to separate the variables. Let us attempt a solution of the form y…x;t†ˆu…x†f…t†. Substitution of this into the partial di/C128erential equation yields two equations: f00…t†‡/C332f…t†ˆ0;xu00…x†‡u0…x†‡…/C332=/C103†u…x†ˆ0; where /C332is the separation constant. The di/C128erential equation for f…t†is ready for integration and the result is f…t†ˆcos…/C33tÿ†, with a phase constant. The di/C128erential equation for u…x†is not in a recognizable form yet. To solve it, first change variables by putting xˆ/C103z2=4;/C119…z†ˆu…x†; then the di/C128erential equation for u…x†becomes Bessel’s equation of order zero: z/C11900…z†‡/C1190…z†‡/C332z/C119…z†ˆ0: Its general solution is /C119…z†ˆA/C740…/C33z†‡BY0…/C33z† or u…x†ˆA/C7402/C33x /C103/C114 ‡BY02/C33x /C103/C114 : 329BESSEL’S EQUATION Figure 7.5. A flexible chain. Since Y0…2/C33 x=/C103/C112 †!ÿ 1 asx!0, we are forced by physics to choose Bˆ0 and then y…x;t†ˆA/C7402/C33x /C103/C114 cos…/C33tÿ†: The upper end of the chain at xˆ/C108is fixed, requiring that /C7402/C33 ‘ /C103/C115/C32! ˆ0: The frequencies of the normal vibrations of the chain are given by 2/C33n ‘ /C103/C115 ˆ n; where nare the roots of /C740. Some values of /C740…x†and/C741…x†are tabulated at the end of this chapter. Generating function for /C74n…x† The function …x;t†ˆe…x=2†…tÿtÿ1†ˆX1 nˆÿ1/C74n…x†tn…7:87† is called the generating function for Bessel functions of the first kind of integral order. It is very useful in obtaining properties of /C74n…x†for integral values of n which can then often be proved for all values of n. To prove Eq. (7.87), let us consider the exponential functions ext=2andeÿxt=2. The Laurent expansions for these two exponential functions about tˆ0 are ext=2ˆX1 kˆ0…xt=2†k k/C33;eÿxt=2ˆX1 mˆ0…ÿxt=2†k m/C33: Multiplying them together, we get ex…tÿtÿ1†=2ˆX1 kˆ0X1 mˆ0…ÿ1†m k/C33m/C33x 2k‡m tkÿm: …7:88† It is easy to recognize that the coecient of the t0term which is made up of those terms with kˆmis just /C740…x†: X1 kˆ0…ÿ1†k 22k…k/C33†2x2kˆ/C740…x†: 330SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Similarly, the coecient of the term tnwhich is made up of those terms for which kÿmˆnis just /C74n…x†: X1 kˆ0…ÿ1†k …k‡n†/C33k/C3322k‡nx2k‡nˆ/C74n…x†: This shows clearly that the coecients in the Laurent expansion (7.88) of the generating function are just the Bessel functions of integral order. Thus we have proved Eq. (7.87). Bessel’s integral representation With the help of the generating function, we can express /C74n…x†in terms of a definite integral with a parameter. To do this, let tˆeiin the generating func- tion, then ex…tÿtÿ1†=2ˆex…eiÿeÿi†=2ˆeixsin ˆcos…xsin†‡isin…xcos†: Substituting this into Eq. (7.87) we obtain cos…xsin†‡isin…xcos†ˆX1 nˆÿ1/C74n…x†…cos‡isin†n ˆX1 ÿ1/C74n…x†cosn‡iX1 ÿ1/C74n…x†sinn: Since /C74ÿn…x†ˆ… ÿ 1†n/C74n…x†;cosnˆcos…ÿn†, and sin nˆÿsin…ÿn†, we have, upon equating the real and imaginary parts of the above equation, cos…xsin†ˆ/C740…x†‡2X1 nˆ1/C742n…x†cos 2n; sin…xsin†ˆ2X1 nˆ1/C742nÿ1…x†sin…2nÿ1†: It is interesting to note that these are the Fourier cosine and sine series of cos…xsin†and sin …xsin†. Multiplying the first equation by cos kand integrat- ing from 0 to , we obtain 1 Z 0coskcos…xsin†dˆ/C74k…x†;ifkˆ0;2;4;... 0; ifkˆ1;3;5;...( : 331BESSEL’S EQUATION Now multiplying the second equation by sin kand integrating from 0 to ,w e obtain 1 Z 0sinksin…xsin†dˆ/C74k…x†;ifkˆ1;3;5;... 0; ifkˆ0;2;4;...( : Adding these two together we obtain Bessel’s integral representation /C74n…x†ˆ1 Z 0cos…nÿxsin†d;nˆpositive integer : …7:89† Recurrence formulas for /C74n…x† Bessel functions of the first kind, /C74n…x†, are the most useful, because they are bounded near the origin. And there exist some useful recurrence formulas between Bessel functions of di/C128erent orders and their derivatives. …1†/C74n‡1…x†ˆ2n x/C74n…x†ÿ/C74nÿ1…x†: …7:90† Proof: Di/C128erentiating both sides of the generating function with respect to t,w e obtain ex…tÿtÿ1†=2x 21‡1 t2 ˆX1 nˆÿ1n/C74n…x†tnÿ1 or x 21‡1 t2X1 nˆÿ1/C74n…x†tnˆX1 nˆÿ1n/C74n…x†tnÿ1: This can be rewritten as x 2X1 nˆÿ1/C74n…x†tn‡x 2X1 nˆÿ1/C74n…x†tnÿ2ˆX1 nˆÿ1n/C74n…x†tnÿ1 or x 2X1 nˆÿ1/C74n…x†tn‡x 2X1 nˆÿ1/C74n‡2…x†tnˆX1 nˆÿ1…n‡1†/C74n‡1…x†tn: Equating coecients of tnon both sides, we obtain x 2/C74n…x†‡x 2/C74n‡2…x†ˆ…n‡1†/C74n…x†: Replacing nbynÿ1, we obtain the required result. …2†x/C740 n…x†ˆn/C74n…x†ÿx/C74n‡1…x†: …7:91† 332SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Proof: /C74n…x†ˆX1 kˆ0…ÿ1†k k/C33ÿ…n‡k‡1†2n‡2kxn‡2k: Di/C128erentiating both sides once, we obtain /C740 n…x†ˆX1 kˆ0…n‡2k†…ÿ1†k k/C33ÿ…n‡k‡1†2n‡2kxn‡2kÿ1; from which we have x/C740 n…x†ˆn/C74n…x†‡xX1 kˆ1…ÿ1†k …kÿ1†/C33ÿ…n‡k‡1†2n‡2kÿ1xn‡2kÿ1: Letting kˆm‡1 in the sum on the right hand side, we obtain x/C740 n…x†ˆn/C74n…x†ÿxX1 mˆ0…ÿ1†m m/C33ÿ…n‡m‡2†2n‡2m‡1xn‡2m‡1 ˆn/C74n…x†ÿx/C74n‡1…x†: …3†x/C740 n…x†ˆÿ n/C74n…x†‡x/C74nÿ1…x†: …7:92† Proof: Di/C128erentiating both sides of the following equation with respect to x xn/C74n…x†ˆX1 kˆ0…ÿ1†k k/C33ÿ…n‡k‡1†2n‡2kx2n‡2k; we have d dxfxn/C74n…x†g ˆ xn/C740 n…x†‡nxnÿ1/C74n…x†; d dxX1 kˆ0…ÿ1†kx2n‡2k 2n‡2kk/C33ÿ…n‡k‡1†ˆX1 kˆ0…ÿ1†kx2n‡2kÿ1 2n‡2kÿ1k/C33ÿ…n‡k† ˆxnX1 kˆ0…ÿ1†kx…nÿ1†‡2k 2…nÿ1†‡2kk/C33ÿ‰…nÿ1†‡k‡1Š ˆxn/C74nÿ1…x†: Equating these two results, we have xn/C740 n…x†‡nxnÿ1/C74n…x†ˆxn/C74nÿ1…x†: 333BESSEL’S EQUATION Canceling out the common factor xnÿ1, we obtained the required result (7.92). …4†/C740 n…x†ˆ‰/C74nÿ1…x†ÿ/C74n‡1…x†Š=2: …7:93† Proof: Adding (7.91) and (7.92) and dividing by 2 x, we obtain the required result (7.93). If we subtract (7.91) from (7.92), /C740 n…x†is eliminated and we obtain x/C74n‡1…x†‡x/C74nÿ1…x†ˆ2n/C74n…x† which is Eq. (7.90). These recurrence formulas (or important identities) are very useful. Here are some illustrative examples. Example 7.2 Show that /C740 0…x†ˆ/C74ÿ1…x†ˆÿ /C741…x†. Solution: From Eq. (7.93), we have /C740 0…x†ˆ‰/C74ÿ1…x†ÿ/C741…x†Š=2; then using the fact that /C74ÿn…x†ˆ… ÿ 1†n/C74n…x†, we obtain the required results. Example 7.3 Show that /C743…x†ˆ8 x2ÿ1 /C741…x†ÿ4 x/C740…x†: Solution: Letting nˆ4 in (7.90), we have /C743…x†ˆ4 x/C742…x†ÿ/C741…x†: Similarly, for /C742…x†we have /C742…x†ˆ2 x/C741…x†ÿ/C740…x†: Substituting this into the expression for /C743…x†, we obtain the required result. Example 7.4FindR t 0x/C740…x†dx. 334SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Solution: Taking derivative of the quantity x/C741…x†with respect to x, we obtain d dxfx/C741…x†g ˆ /C741…x†‡x/C740 1…x†: Then using Eq. (7.92) with nˆ1,x/C740 1…x†ˆÿ /C741…x†‡x/C740…x†, we find d dxfx/C741…x†g ˆ /C741…x†‡x/C740 1…x†ˆx/C740…x†; thus, Zt 0x/C740…x†dxˆx/C741…x†jt 0ˆt/C741…t†: /C65pproximations to the Bessel functions For very large or very small values of xwe might be able to make some approxi- mations to the Bessel functions of the first kind /C74n…x†. By a rough argument, we can see that the Bessel functions behave something like a damped cosine function when the value of xis very large. To see this, let us go back to Bessel’s equation (7.71) x2y00‡xy0‡…x2ÿ 2†yˆ0 and rewrite it as y00‡1 xy0‡1ÿ 2 x2/C32! yˆ0: Ifxis very large, let us drop the term 2=x2and then the di/C128erential equation reduces to y00‡1 xy0‡yˆ0: Letuˆyx1=2, then u0ˆy0x1=2‡1 2xÿ1=2y, and u00ˆy00x1=2‡xÿ1=2y0ÿ14xÿ3=2y. From u00we have y00‡1 xy0ˆxÿ1=2u00‡1 4x2y: Adding yon both sides, we obtain y00‡1 xy0‡yˆ0ˆxÿ1=2u00‡1 4x2y‡y; xÿ1=2u00‡1 4x2y‡yˆ0 335BESSEL’S EQUATION or u00‡1 4x2‡1 x1=2yˆu00‡1 4x2‡1 uˆ0; the solution of which is uˆAcosx‡Bsinx: Thus the approximate solution to Bessel’s equation for very large values of xis yˆxÿ1=2…Acosx‡Bsinx†ˆCxÿ1=2cos…x‡/C12†: A more rigorous argument leads to the following asymptotic formula /C74n…x†/C252 x1=2 cosxÿ 4ÿn 2 : …7:94† For very small values of x(that is, near 0), by examining the solution itself and dropping all terms after the first, we find /C74n…x†/C25xn 2nÿ…n‡1†: …7:95† Orthogonality of Bessel functions Bessel functions enjoy a property which is called orthogonality and is of general importance in mathematical physics. If and/C22are two di/C128erent constants, we can show that under certain conditions Z1 0x/C74n…x†/C74n…/C22x†dxˆ0: Let us see what these conditions are. First, we can show that Z1 0x/C74n…x†/C74n…/C22x†dxˆ/C22/C74n…†/C740 n…/C22†ÿ/C74n…/C22†/C740 n…† 2ÿ/C222: …7:96† To show this, let us go back to Bessel’s equation (7.71) and change the indepen-dent variable to x, where is a constant, then the resulting equation is x 2y00‡xy0‡…2x2ÿn2†yˆ0 and its general solution is /C74n…x†. Now suppose we have two such equations, one fory1with constant , and one for y2with constant /C22: x2y00 1‡xy0 1‡…2x2ÿn2†y1ˆ0;x2y00 2‡xy0 2‡…/C222x2ÿn2†y2ˆ0: Now multiplying the first equation by y2, the second by y1and subtracting, we get x2‰y2y00 1ÿy1y00 2Їx‰y2y0 1ÿy1y0 2Šˆ…/C222ÿ2†x2y1y2: 336SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS Dividing by xwe obtain xd dx‰y2y0 1ÿy1y0 2Ї‰y2y0 1ÿy1y0 2Šˆ…/C222ÿ2†xy1y2 or d dxfx‰y2y0 1ÿy1y0 2Šg ˆ … /C222ÿ2†xy1y2 and then integration gives …/C222ÿ2†Z xy1y2dxˆx‰y2y0 1ÿy1y0 2Š; where we have omitted the constant of integration. Now y1ˆ/C74n…x†;y2ˆ/C74n…x†, and if 6ˆ/C22we then have Z x/C74n…x†/C74n…/C22x†dxˆx‰/C74n…/C22x†/C740 n…x†ÿ/C22/C74n…x†/C740 n…/C22x† /C222ÿ2: Thus Z1 0x/C74n…x†/C74n…/C22x†dxˆ/C22/C74n…†/C740 n…/C22†ÿ/C74n…/C22†/C740 n…† 2ÿ/C222q:e:d: Now letting /C22!and using L’Hospital’s rule, we obtain Z1 0x/C742 n…x†dxˆlim /C22!/C740 n…/C22†/C740 n…†ÿ/C74n…†/C740 n…/C22†ÿ/C22/C74n…†/C7400 n…/C22† 2/C22 ˆ/C740 n2…†ÿ/C74n…†/C740 n…†ÿ/C74n…†/C7400 n…† 2: But 2/C7400 n…†‡/C740 n…†‡…2ÿn2†/C74n…†ˆ0: Solving for /C7400 n…†and substituting, we obtain Z1 0x/C742 n…x†dxˆ1 2/C740 n2…†‡ 1ÿn2 2/C32! /C742 n…x†"# : …7:97† Furthermore, if and /C22are any two di/C128erent roots of the equation R/C74n…x†‡Sx/C740 n…x†ˆ0, where /C82andSare constant, we then have R/C74n…†‡S/C740 n…†ˆ0;R/C74n…/C22†‡S/C22/C740 n…/C22†ˆ0; from these two equations we find, if R6ˆ0;S6ˆ0, /C22/C74n…†/C740 n…/C22†ÿ/C74n…/C22†/C740 n…†ˆ0 337BESSEL’S EQUATION and then from Eq. (7.96) we obtain Z1 0x/C74n…x†/C74n…/C22x†dxˆ0: …7:98† Thus, the two functionsxp/C74n…x†andxp/C74 n…/C22x†are orthogonal in (0, 1). We can also say that the two functions /C74n…x†and/C74n…/C22x†are orthogonal with respect to the weighted function x. Eq. (7.98) is also easily proved if Rˆ0 and S6ˆ0, or R6ˆ0 but Sˆ0. In this case, and/C22can be any two di/C128erent roots of /C74n…x†ˆ0o r/C740 n…x†ˆ0. /C83pherical /C66essel functions In physics we often meet the following equation d drr2dR dr ‡‰k2r2ÿ/C108…/C108‡1†ŠRˆ0; …/C108ˆ0;1;2;...†: …7:99† In fact, this is the radial equation of the wave and the Helmholtz partial di/C128er- ential equation in the spherical coordinate system (see Problem 7.22). If we let xˆkrandy…x†ˆR…r†, then Eq. (7.99) becomes x2y00‡2xy0‡‰x2ÿ/C108…/C108‡1†Šyˆ0 …lˆ0;1;2;...†; …7:100† where y0ˆdy=dx. This equation almost matches Bessel’s equation (7.71). Let us make the further substitution y…x†ˆ/C119…x†=xp; then we obtain x2/C11900‡x/C1190‡‰x2ÿ…/C108‡1 2†Š/C119ˆ0 …/C108ˆ0;1;2;...†: …7:101† The reader should recognize this equation as Bessel’s equation of order /C108‡12.I t follows that the solutions of Eq. (7.100) can be written in the form y…x†ˆA/C74/C108‡1=2…x†xp ‡B/C74ÿ/C108ÿ1=2…x†xp : This leads us to define spherical Bessel functions j /C108…x†ˆC/C74/C108‡/C69…x†=xp. The factor /C67is usually chosen to be =2/C112 for a reason to be explained later: j /C108…x†ˆ =2x/C112 /C74/C108‡/C69…x†: …7:102† Similarly, we can define n/C108…x†ˆ =2x/C112 /C78 /C108‡/C69…x†: 338SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS We can express j/C108…x†in terms of j0…x†. To do this, let us go back to /C74n…x†and we find that d dxfxÿn/C74n…x†g ˆ ÿ xÿn/C74n‡1…x†;or/C74n‡1…x†ˆÿ xnd dxfxÿn/C74n…x†g: The proof is simple and straightforward: d dxfxÿn/C74n…x†g ˆd dxX1 kˆ0…ÿ1†kx2k 2n‡2kk/C33ÿ…n‡k‡1† ˆxÿnX1 kˆ0…ÿ1†kxn‡2kÿ1 2n‡2kÿ1…kÿ1†/C33ÿ…n‡k‡1† ˆxÿnX1 kˆ0…ÿ1†k‡1xn‡2k‡1 2n‡2k‡1k/C33ÿ‰…n‡k‡2gˆÿxÿn/C74n‡1…x†: Now if we set nˆ/C108‡1 2and divide by x/C108‡3=2, we obtain /C74/C108‡3=2…x† x/C108‡3=2ˆÿ1 xd dx/C74/C108‡1=2…x† x/C108‡1=2 orj/C108‡1…x† x/C108‡1ˆÿ1 xd dxj/C108…x† x/C108 : Starting with /C108ˆ0 and applying this formula ltimes, we obtain j/C108…x†ˆx/C108ÿ1 xd dx/C108 j0…x†… /C108ˆ1;2;3;...†: …7:103† Once j0…x†has been chosen, all j/C108…x†are uniquely determined by Eq. (7.103). Now let us go back to Eq. (7.102) and see why we chose the constant factor /C67to be =2/C112 . If we set /C108ˆ0 in Eq. (7.101), the resulting equation is xy00‡2y0‡xyˆ0: Solving this equation by the power series method, the reader will find that func- tions sin ( x†=xand cos ( x†=xare among the solutions. It is customary to define j0…x†ˆsin…x†=x: Now by using Eq. (7.76), we find /C741=2…x†ˆX1 kˆ0…ÿ1†k…x=2†1=2‡2k k/C33ÿ…k‡3=2† ˆ…x=2†1=2 …1=2†p 1ÿx2 3/C33‡x4 5/C33ÿ/C32! ˆ…x=2†1=2 …1=2†psinx xˆ 2 x/C114 sinx: Comparing this with j0…x†shows that j0…x†ˆ =2x/C112 /C741=2…x†, and this explains the factor =2/C112 chosen earlier. 339SPHERICAL BESSEL FUNCTIONS /C83turm/C177Liou/C118ille s/C121stems A boundary-value problem having the form d dxr…x†dy dx ‡‰/C113…x†‡/C112…x†Šyˆ0; axb …7:104† and satisfying boundary conditions of the form k1y…a†‡k2y0…a†ˆ0; /C1081y…b†‡/C1082y0…b†ˆ0 …7:104a† is called a Sturm–Liouville boundary-value problem; Eq. (7.104) is known as the Sturm–Liouville equation. Legendre’s equation, Bessel’s equation and many other important equations can be written in the form of (7.104). Legendre’s equation (7.1) can be written as ‰…1ÿx2†y0Š0‡yˆ0; ˆ/C23…/C23‡1†; we can then see it is a Sturm–Liouville equation with rˆ1ÿx2;/C113ˆ0 and /C112ˆ1. Then, how do Bessel functions fit into the Sturm–Liouville framework/C63 /C74…s† satisfies the Bessel equation (7.71) s2/C127/C74n‡s_/C74n‡…s2ÿn2†/C74nˆ0; _/C74nˆd/C74n=ds: …7:71a† We assume nis a positive integer and setting sˆx, with a non-zero constant, we have ds dxˆ; _/C74nˆd/C74n dxdx dsˆ1 d/C74n dx;/C127/C74nˆd dx1 d/C74n dxdx dsˆ1 2d2/C74n dx2 and Eq. (7.71a) becomes x2/C7400 n…x†‡x/C740 n…x†‡…2x2ÿn2†/C74n…x†ˆ0;/C740 nˆd/C74n=dx or x/C7400 n…x†‡/C740 n…x†‡…2xÿn2=x†/C74n…x†ˆ0; which can be written as ‰x/C740 n…x†Š0‡ÿn2 x‡2x/C32! /C74n…x†ˆ0: It is easy to see that for each fixed nthis is a Sturm–Liouville equation (7.104), with r…x†ˆx,/C113…x†ˆÿ n2=x;/C112…x†ˆx, and with the parameter now written as 2. For the Sturm–Liouville system (7.104) and (7.104a), a non-trivial solution exists in general only for a particular set of values of the parameter . These values are called the eigenvalues of the system. If r…x†and/C113…x†are real, the eigenvalues are real. The corresponding solutions are called eigenfunctions of 340SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS the system. In general there is one eigenfunction to each eigenvalue. This is the non-degenerate case. In the degenerate case, more than one eigenfunction maycorrespond to the same eigenvalue. The eigenfunctions form an orthogonal set with respect to the density function /C112…x†which is generally 0.Thus by suitable normalization the set of functions can be made an orthonormal set with respect to/C112…x†inaxb. We now proceed to prove these two general claims. Propert/C121 /C49Ifr…x†and/C113…x†are real, the eigenvalues of a Sturm–Liouville system are real. We start with the Sturm–Liouville equation (7.104) and the boundary condi- tions (7.104a): d dxr…x†dy dx ‡‰/C113…x†‡/C112…x†Šyˆ0; axb; k1y…a†‡k2y0…a†ˆ0;/C1081y…b†‡/C1082y0…b†ˆ0; and assume that r…x†;/C113…x†;/C112…x†;k1;k2;/C1081,a n d /C1082are all real, but andymay be complex. Now take the complex conjugates d dxr…x†dy dx ‡‰/C113…x†‡ /C112…x†Šyˆ0; …7:105† k1y…a†‡k2y0…a†ˆ0;/C1081y…b†‡/C1082y0…b†ˆ0; …7:105a† where yand are the complex conjugates of yand, respectively. Multiplying (7.104) by y, (7.105) by y, and subtracting, we obtain after simplifying d dxr…x†…yy0ÿyy0†/C2/C3 ˆ…ÿ†/C112…x†yy: Integrating from atob, and using the boundary conditions (7.104a) and (7.105a), we then obtain …ÿ†Zb a/C112…x†y0/C12/C12/C12/C122dxˆr…x†…yy0ÿyy0†jb aˆ0: Since /C112…x†0i n axb, the integral on the left is positive and therefore ˆ, that is, is real. Propert/C121 /C50 The eigenfunctions corresponding to two di/C128erent eigenvalues are orthogonal with respect to /C112…x†inaxb. 341STURM–LIOUVILLE SYSTEMS Ify1andy2are eigenfunctions corresponding to the two di/C128erent eigenvalues 1;2, respectively, d dxr…x†dy1 dx ‡‰/C113…x†‡1/C112…x†Šy1ˆ0; axb; …7:106† k1y1…a†‡k2y0 1…a†ˆ0;/C1081y1…b†‡/C1082y0 1…b†ˆ0; …7:106a† d dxr…x†dy2 dx ‡‰/C113…x†‡2/C112…x†Šy2ˆ0; axb; …7:107† k1y2…a†‡k2y0 2…a†ˆ0; /C1081y2…b†‡/C1082y0 2…b†ˆ0: …7:107a† Multiplying (7.106) by y2and (7.107) by y1, then subtracting, we obtain d dxr…x†…y1y0 2ÿy2y0 1†/C2/C3 ˆ…ÿ†/C112…x†y1y2: Integrating from atob, and using (7.106a) and (7.107a), we obtain …1ÿ2†Zb a/C112…x†y1y2dxˆr…x†…y1y0 2ÿy2y0 1†jb aˆ0: Since 16ˆ2we have the required result; that is, Zb a/C112…x†y1y2dxˆ0: We can normalize these eigenfunctions to make them an orthonormal set, and so we can expand a given function in a series of these orthonormal eigenfunctions. We have shown that Legendre’s equation is a Sturm–Liouville equation with r…x†ˆ1ÿx;/C113ˆ0 and /C112ˆ1. Since rˆ0 when xˆ1, no boundary conditions are needed to form a Sturm–Liouville problem on the interval ÿ1x1. The numbers nˆn…n‡1†are eigenvalues with nˆ0;1;2;3;.... The corresponding eigenfunctions are ynˆPn…x†. Property 2 tells us that Z1 ÿ1Pn…x†Pm…x†dxˆ0 n6ˆm: For Bessel functions we saw that ‰x/C740 n…x†Š0‡ÿn2 x‡2x/C32! /C74n…x†ˆ0 is a Sturm–Liouville equation (7.104), with r…x†ˆx;/C113…x†ˆÿ n2=x;/C112…x†ˆx, and with the parameter now written as 2. Typically, we want to solve this equation 342SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS on an interval 0 xbsubject to /C74n…b†ˆ0: which limits the selection of . Property 2 then tells us that Zb 0x/C74n…kx†/C74n…/C108x†dxˆ0; k6ˆ/C108: Problems 7.1 Using Eq. (7.11), show that Pn…ÿx†ˆ… ÿ 1†nPn…x†and Pn0…ÿx†ˆ …ÿ1†n‡1P0 n…x†: 7.2 Find P0…x†;P1…x†;P2…x†;P3…x†, and P4…x†from Rodrigues’ formula (7.12). Compare your results with Eq. (7.11). 7.3 Establish the recurrence formula (7.16b) by manipulating Rodrigues’ formula. 7.4 Prove that P0 5…x†ˆ9P4…x†‡5P2…x†‡P0…x†. Hint: Use the recurrence relation (7.16d). 7.5 Let Pand/C81be two points in space (Fig. 7.6). Using Eq. (7.14), show that 1 rˆ1 r2 1‡r22ÿ2r1r2cosq ˆ1 r2P0‡P1…cos†r1 r2‡P2…cos†r1 r22 ‡"# : 7.6 What is Pn…1†/C63What is Pn…ÿ1†/C63 7.7 Obtain the associated Legendre functions: …a†P1 2…x†;…b†P23…x†;…c†P32…x†: 7.8 Verify that P2 3…x†is a solution of Legendre’s associated equation (7.25) for mˆ2,nˆ3. 7.9 Verify the orthogonality conditions (7.31) for the functions P1 2…x†andP13…x†. 343PROBLEMS Figure 7.6. 7.10 Verify Eq. (7.37) for the function P1 2…x†: 7.11 Show that dnÿm dxnÿm…x2ÿ1†nˆ…nÿm†/C33 …n‡m†/C33…x2ÿ1†mdn‡m dxn‡m…x2ÿ1†m Hint: Write …x2ÿ1†nˆ…xÿ1†n…x‡1†nand find the derivatives by Leibnitz’s rule. 7.12 Use the generating function for the Hermite polynomials to find: (a)H0…x†; (b) H1…x†;(c)H2…x†;( d ) H3…x†. 7.13 Verify that the generating function satisfies the identity /C642 /C64x2ÿ2x/C64 /C64x‡2t/C64 /C64tˆ0: Show that the functions Hn…x†in Eq. (7.47) satisfy Eq. (7.38). 7.14 Given the di/C128erential equation y00‡…/C34ÿx2†yˆ0, find the possible values of/C34(eigenvalues) such that the solution y…x†of the given di/C128erential equa- tion tends to zero as x! 1 . For these values of /C34, find the eigenfunctions y…x†. 7.15 In Eq. (7.58), write the series for the exponential and collect powers of zto verify the first few terms of the series. Verify the identity x/C642 /C64x2‡…1ÿx†/C64 /C64x‡z/C64 /C64zˆ0: Substituting the series (7.58) into this identity, show that the functions Ln…x† in Eq. (7.58) satisfy Laguerre’s equation. 7.16 Show that /C740…x†ˆ1ÿx2 22…1/C33†2‡x4 24…2/C33†2ÿx6 26…3/C33†2‡ÿ ; /C741…x†ˆx 2ÿx3 231/C332/C33‡x5 252/C333/C33ÿx7 273/C334/C33‡ÿ : 7.17 Show that /C741=2…x†ˆ2 x1=2 sinx;/C74ÿ1=2…x†ˆ2 x1=2 cosx: 7.18 If nis a positive integer, show that the formal expression for /C74ÿn…x†gives /C74ÿn…x†ˆ… ÿ 1†n/C74n…x†. 7.19 Find the general solution to the modified Bessel’s equation x2y00‡xy0‡…x2s2ÿ 2†yˆ0 which di/C128ers from Bessel’s equation only in that sxtakes the place of x. 344SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS (Hint: Reduce the given equation to Bessel’s equation first.) 7.20 The lengthening simple pendulum: Consider a small mass msuspended by a string of length l. If its length is increased at a steady rate ras it swings back and forth freely in a vertical plane, find the equation of motion and the solution for small oscillations. 7.21 Evaluate the integrals: …a†Z xn/C74nÿ1…x†dx; …b†Z xÿn/C74n‡1…x†dx; …c†Z xÿ1/C741…x†dx: 7.22 In quantum mechanics, the three-dimensional Schro /C200dinger equation is ip/C64/C32…r;t† /C64tˆÿp2 2m/C1142/C32…r;t†‡/C86/C32…r;t†; iˆ  ÿ1p ;pˆ/C104=2: (a) When the potential /C86is independent of time, we can write /C32…r;t†ˆ u…r†T…t†. Show that in this case the Schro /C200dinger equation reduces to ÿp2 2m/C1142u…r†‡/C86u…r†ˆ/C69u…r†; a time-independent equation along with T…t†ˆeÿi/C69t=p, where Eis a separation constant. (b) Show that, in spherical coordinates, the time-independent Schro /C200dinger equation takes the form ÿp2 2m1 r2/C64 /C64rr2/C64u /C64r ‡1 r2sin/C64 /C64sin/C64u /C64 ‡1 r2sin2/C642u /C64/C30"# ‡/C86…r†uˆ/C69u; then use separation of variables, u…r; ;/C30†ˆR…r†Y…; /C30†, to split it into two equations, with as a new separation constant: ÿp2 2m1 r2d drr2dR dr ‡/C86‡ r2 Rˆ/C69R; ÿp2 2m1 sin/C64 /C64sin/C64Y /C64 ÿp2 2m1 sin2/C642Y /C64/C302ˆ Y: It is straightforward to see that the radial equation is in the form of Eq. (7.99). Continuing the separation process by putting Y…; /C30†ˆ/C2…†…†, the angular equation can be separated further into two equations, with/C12as separation constant: ÿp2 2m1 d2 d/C302ˆ/C12; ÿp2 2msind dsind/C2 d ÿ sin2/C2‡/C12/C2ˆ0: 345PROBLEMS The first equation is ready for integration. Do you recognize the second equation in as Legendre’s equation/C63 (Compare it with Eq. (7.30).) If you are unsure, try to simplify it by putting /C13ˆ2m =p; /C22ˆ…2m/C12=p†1=2, and you will obtain sind dsind/C2 d ‡…/C13sin2ÿ/C222†/C2ˆ0 or 1 sind dsind/C2 d ‡ /C13ÿ/C222 sin2 /C2ˆ0; which more closely resembles Eq. (7.30). 7.23 Consider the di/C128erential equation y00‡R…x†y0‡‰/C81…x†‡P…x†Šyˆ0: Show that it can be put into the form of the Sturm–Liouville equation (7.104) with r…x†ˆeR R…x†dx;/C113…x†ˆ/C81…x†eR R…x†dx;and /C112…x†ˆP…x†eR R…x†dx: 7.24. ( a) Show that the system y00‡yˆ0;y…0†ˆ0;y…1†ˆ0 is a Sturm– Liouville system. (b) Find the eigenvalues and eigenfunctions of the system. (c) Prove that the eigenfunctions are orthogonal on the interval 0 x1. (d) Find the corresponding set of normalized eigenfunctions, and expand the function f…x†ˆ1 in a series of these orthonormal functions. 346SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS 8 The calculus of variations The calculus of variations, in its present form, provides a powerful method for the treatment of variational principles in physics and has become increasingly impor- tant in the development of modern physics. It is originated as a study of certain extremum (maximum and minimum) problems not treatable by elementary calculus. To see this more precisely let us consider the following integral whose integrand is a function of x,y, and of the first derivative y0…x†ˆdy=dx: IˆZx2 x1fy…x†;y0…x†;x/C8/C9 dx; …8:1† where the semicolon in fseparates the independent variable xfrom the dependent variable y…x†and its derivative y0…x†. For what function y…x†is the value of the integral Ia maximum or a minimum/C63 This is the basic problem of the calculus of variations. The quantity fdepends on the functional form of the dependent variable y…x† and is called the functional which is considered as given, the limits of integrationare also given. It is also understood that yˆy 1atxˆx1,yˆy2atxˆx2.I n contrast with the simple extreme-value problem of di/C128erential calculus, the func-tiony…x†is not known here, but is to be varied until an extreme value of the integral Iis found. By this we mean that if y…x†is a curve which gives to Ia minimum value, then any neighboring curve will make Iincrease. We can make the definition of a neighboring curve clear by giving y…x†a parametric representation: y…/C34;x†ˆy…0;x†‡/C34/C17…x†; …8:2† where /C17…x†is an arbitrary function which has a continuous first derivative and /C34is a small arbitrary parameter. In order for the curve (8.2) to pass through …x 1;y1† and…x2;y2†, we require that /C17…x1†ˆ/C17…x2†ˆ0 (see Fig. 8.1). Now the integral I 347 also becomes a function of the parameter /C34 I…/C34†ˆZx2 x1ffy…/C34;x†;y0…/C34;x†;xgdx: …8:3† We then require that y…x†ˆy…0;x†makes the integral Ian extreme, that is, the integral I…/C34†has an extreme value for /C34ˆ0: I…/C34†ˆZx2 x1ffy…/C34;x†;y0…/C34;x†;xgdxˆextremum for /C34ˆ0: This gives us a very simple method of determining the extreme value of the integral I. The necessary condition is dI d/C34/C12/C12/C12/C12 /C34ˆ0ˆ0 …8:4† for all functions /C17…x†. The sucient conditions are quite involved and we shall not pursue them. The interested reader is referred to mathematical texts on the calculus of variations. The problem of the extreme-value of an integral occurs very often in geometry and physics. The simplest example is provided by the problem of determining the shortest curve (or distance) between two given points. In a plane, this is the straight line. But if the two given points lie on a given arbitrary surface, then the analytic equation of this curve, which is called a geodesic, is found by solution of the above extreme-value problem. /C84he /C69uler/C177Lagrange equation In order to find the required curve y…x†we carry out the indicated di/C128erentiation in the extremum condition (8.4): 348THE CALCULUS OF VARIATIONS Figure 8.1. /C64I /C64/C34ˆ/C64 /C64/C34Zx2 x1ffy…/C34;x†;y0…/C34;x†;xgdx ˆZx2 x1/C64f /C64y/C64y /C64/C34‡/C64f /C64y0/C64y0 /C64/C34 dx; …8:5† where we have employed the fact that the limits of integration are fixed, so the di/C128erential operation a/C128ects only the integrand. From Eq. (8.2) we have /C64y /C64/C34ˆ/C17…x†and/C64y0 /C64/C34ˆd/C17 dx: Substituting these into Eq. (8.5) we obtain /C64I /C64/C34ˆZx2 x1/C64f /C64y/C17…x†‡/C64f /C64y0d/C17 dx dx: …8:6† Using integration by parts, the second term on the right hand side becomes Zx2 x1/C64f /C64y0d/C17 dxdxˆ/C64f /C64y0/C17…x†/C12/C12/C12/C12x2 x1ÿZx2 x1d dx/C64f /C64y0 /C17…x†dx: The integrated term on the right hand side vanishes because /C17…x1†ˆ/C17…x2†ˆ0 and Eq. (8.6) becomes /C64I /C64/C34ˆZx2 x1/C64f /C64y/C64y /C64/C34ÿd dx/C64f /C64y0/C64y /C64/C34 dx ˆZx2 x1/C64f /C64yÿd dx/C64f /C64y0  /C17…x†dx: …8:7† Note that /C64f=/C64yand /C64f=/C64y0are still functions of /C34. However, when /C34ˆ0;y…/C34;x†ˆy…x†and the dependence on /C34disappears. Then …/C64I=/C64/C34†j/C34ˆ0vanishes, and since /C17…x†is an arbitrary function, the inte- grand in Eq. (8.7) must vanish for /C34ˆ0: d dx/C64f /C64y0ÿ/C64f /C64yˆ0: …8:8† Eq. (8.8) is known as the Euler–Lagrange equation; it is a necessary but not sucient condition that the integral Ihave an extreme value. Thus, the solution of the Euler–Lagrange equation may not yield the minimizing curve. Ordinarilywe must verify whether or not this solution yields the curve that actually mini- mizes the integral, but frequently physical or geometrical considerations enable us to tell whether the curve so obtained makes the integral a minimum or a max- imum. The Euler–Lagrange equation can be written in the form (Problem 8.2) d dxfÿy0/C64f /C64y0 ÿ/C64f /C64xˆ0: …8:8a† 349THE EULER–LAGRANGE EQUATION This is often called the second form of the Euler–Lagrange equation. If fdoes not involve xexplicitly, it can be integrated to yield fÿy0/C64f /C64y0ˆc; …8:8b† where cis an integration constant. The Euler–Lagrange equation can be extended to the case in which fis a functional of several dependent variables: fˆfy 1…x†;y0 1…x†;y2…x†;y0 2…x†;...;x/C8/C9 : Then, in analogy with Eq. (8.2), we now have yi…/C34;x†ˆyi…0;x†‡/C34/C17i…x†; iˆ1;2;...;n: The development proceeds in an exactly analogous manner, with the result /C64I /C64/C34ˆZx2 x1/C64f /C64yiÿd dx/C64f /C64yi0  /C17i…x†dx: Since the individual variations, that is, the /C17i…x†, are all independent, the vanish- ing of the above equation when evaluated at /C34ˆ0 requires the separate vanishing of each expression in the brackets: d dx/C64f /C64y0 iÿ/C64f /C64yiˆ0; iˆ1;2;...;n: …8:9† Example 8.1 The brachistochrone problem: Historically, the brachistochrone problem was the first to be treated by the method of the calculus of variations (first solved by Johann Bernoulli in 1696). As shown in Fig. 8.2, a particle is constrained to move in a gravitational field starting at rest from some point P1to some lower 350THE CALCULUS OF VARIATIONS Figure 8.2 point P2. Find the shape of the path such that the particle goes from P1toP2in the least time. (The word brachistochrone was derived from the Greek brachistos (shortest) and chronos (time).) Solution: If O andPare not very far apart, the gravitational field is constant, and if we ignore the possibility of friction, then the total energy of the particle is conserved: 0‡m/C103y 1ˆ1 2mds dt2 ‡m/C103…y1ÿy†; where the left hand side is the sum of the kinetic energy and the potential energy of the particle at point P1, and the right hand side refers to point P…x;y†. Solving fords=dt: ds=dtˆ 2/C103y/C112 : Thus the time required for the particle to move from P1toP2is tˆZP2 P1dtˆZP2 P1ds2/C103yp : The line element dscan be expressed as dsˆ dx2‡dy2q ˆ1‡y 02q dx; y0ˆdy=dx; thus, we have tˆZP2 P1dtˆZP2 P1ds2/C103yp ˆ12/C103pZ x2 0 1‡y02/C112 yp dx: We now apply the Euler–Lagrange equation to find the shape of the path for the particle to go from P1toP2in the least time. The constant does not a/C128ect the final equation and the functional fmay be identified as fˆ 1‡y02q =yp; which does not involve xexplicitly. Using Problem 8.2(b), we find fÿy0/C64f /C64y0ˆ 1‡y02/C112 yp ÿy0 y0  1‡y02/C112 yp"# ˆc; which simplifies to  1‡y02qypˆ1=c: 351THE EULER–LAGRANGE EQUATION Letting 1 =cˆapand solving for y0gives y0ˆdy dxˆaÿy y/C114 ; and solving for dxand integrating we obtain Z dxˆZy aÿy/C114 dy: We then let yˆasin2ˆa 2…1ÿcos 2 † which leads to xˆ2aZ sin2dˆaZ …1ÿcos 2 †dˆa2…2ÿsin 2 †‡k: Thus the parametric equation of the path is given by xˆb…1ÿcos/C30†; yˆb…/C30ÿsin/C30†‡k; where bˆa=2;/C30ˆ2. The path passes through the origin so we have kˆ0 and xˆb…1ÿcos/C30†; yˆb…/C30ÿsin/C30†: The constant bis determined from the condition that the particle passes through P 2…x2;y2†: The required path is a cycloid and is the path of a fixed point P0on a circle of radius bas it rolls along the x-axis (Fig. 8.3). A line that represents the shortest path between any two points on some surface is called a geodesic. On a flat surface, the geodesic is a straight line. It is easy to show that, on a sphere, the geodesic is a great circle; we leave this as an exercise for the reader (Problem 8.3). 352THE CALCULUS OF VARIATIONS Figure 8.3. /C86ariational problems /C119ith constraints In certain problems we seek a minimum or maximum value of the integral (8.1) IˆZx2 x1fy…x†;y0…x†;x/C8/C9 dx …8:1† subject to the condition that another integral /C74ˆZx2 x1/C103y…x†;y0…x†;x/C8/C9 dx …8:10† has a known constant value. A simple problem of this sort is the problem of determining the curve of a given perimeter which encloses the largest area, or finding the shape of a chain of fixed length which minimizes the potential energy. In this case we can use the method of Lagrange multipliers which is based on the following theorem: The problem of the stationary value of /C70(x/C44 y) subject to the con/C45 dition /C71…x;y†ˆc/C111nst .is e/C113uivalent to the problem of stationary values/C44 without constraint/C44 of /C70‡/C71for some constant /C44 pro/C45 vided either /C64/C71=/C64xor/C64/C71=/C64ydoes not vanish at the critical point. The constant is called a Lagrange multiplier and the method is known as the method of Lagrange multipliers. To see the ideas behind this theorem, let us assume that /C71…x;y†ˆ0 defines yas a unique function of x, say, yˆ/C103…x†, having a continuous derivative /C1030…x†. Then /C70…x;y†ˆ/C70‰x;/C103…x†Š and its maximum or minimum can be found by setting the derivative with respect toxequal to zero: /C64/C70 /C64x‡/C64/C70 /C64ydy dxˆ0o r /C70x‡/C70y/C1030…x†ˆ0: …8:11† We also have /C71‰x;/C103…x†Š ˆ0; from which we find /C64/C71 /C64x‡/C64/C71 /C64ydy dxˆ0o r /C71x‡/C71y/C1030…x†ˆ0: …8:12† Eliminating /C1030…x†between Eq. (8.11) and Eq. (8.12) we obtain /C70xÿ/C70y=/C71yÿ /C71xˆ0; …8:13† 353VARIATIONAL PROBLEMS WITH CONSTRAINTS provided /C71yˆ/C64/C71=/C64y6ˆ0. Defining ˆÿ/C70y=/C71yor /C70y‡/C71yˆ/C64/C70 /C64y‡/C64/C71 /C64yˆ0; …8:14† Eq. (8.13) becomes /C70x‡/C71xˆ/C64/C70 /C64x‡/C64/C71 /C64xˆ0: …8:15† If we define H…x;y†ˆ/C70…x;y†‡/C71…x;y†; then Eqs. (8.14) and (8.15) become /C64H…x;y†=/C64xˆ0; H…x;y†=/C64yˆ0; and this is the basic idea behind the method of Lagrange multipliers. It is natural to attempt to solve the problem Iˆminimum subject to the con- dition /C74ˆconstant by the method of Lagrange multipliers. We construct the integral I‡/C74ˆZx2 x1‰/C70…y;y0;x†‡/C71…y;y0;x†Šdx and consider its free extremum. This implies that the function y…x†that makes the value of the integral an extremum must satisfy the equation d dx/C64…/C70‡/C71† /C64y0/C64…/C70‡/C71† /C64yˆ0 …8:16† or d dx/C64/C70 /C64y0 ÿ/C64/C70 /C64y ‡d dx/C64/C71 /C64y0 ÿ/C64/C71 /C64y ˆ0: …8:16a† Example 8.2 Isoperimetric problem: Find that curve /C67having the given perimeter lthat encloses the largest area. Solution: The area bounded by /C67can be expressed as Aˆ1 2Z C…xdyÿydx†ˆ12Z C…xy0ÿy†dx and the length of the curve /C67is sˆZ C 1‡y02q dxˆ/C108: 354THE CALCULUS OF VARIATIONS Then the function /C72is HˆZ C‰1 2…xy0ÿy†‡ 1‡y02/C112 Šdx and the Euler–Lagrange equation gives d dx1 2x‡y0  1‡y02/C112/C32! ‡1 2ˆ0 or y0  1‡y02/C112 ˆÿx‡c1: Solving for y0, we get y0ˆdy dxˆxÿc1 2ÿ…xÿc1†2q ; which on integrating gives yÿc2ˆ 2ÿ…xÿc1†2q or …xÿc1†2‡…yÿc2†2ˆ2;a circle : /C72amilton/C39s principle and Lagrange/C39s equation of motion One of the most important applications of the calculus of variations is in classicalmechanics. In this case, the functional fin Eq. (8.1) is taken to be the Lagrangian Lof a dynamical system. For a conservative system, the Lagrangian Lis defined as the di/C128erence of kinetic and potential energies of the system: LˆTÿ/C86; where time tis the independent variable and the generalized coordinates /C113 i…t†are the dependent variables. What do we mean by generalized coordinates/C63 Any convenient set of parameters or quantities that can be used to specify the config- uration (or state) of the system can be assumed to be generalized coordinates; therefore they need not be geometrical quantities, such as distances or angles. In suitable circumstances, for example, they could be electric currents. Eq. (8.1) now takes the form that is known as the action (or the action integral) IˆZt2 t1L/C113 i…t†;_/C113i…t†;t …† dt; _/C113ˆd/C113=dt …8:17† 355HAMILTON’S PRINCIPLE and Eq. (8.4) becomes Iˆ/C64I /C64/C34/C12/C12/C12/C12 /C34ˆ0d/C34ˆZt2 t1L/C113 i…t†; _/C113i…t†;t …† dtˆ0; …8:18† where /C113i…t†, and hence _/C113i…t†, is to be varied subject to /C113i…t1†ˆ/C113i…t2†ˆ0. Equation (8.18) is a mathematical statement of Hamilton’s principle of classical mechanics. In this variational approach to mechanics, the Lagrangian Lis given, and/C113i…t†taken on the prescribed values at t1andt2, but may be arbitrarily varied for values of tbetween t1andt2. In words, Hamilton’s principle states that for a conservative dynamical system, the motion of the system from its position in configuration space at time t1to its position at time t2follows a path for which the action integral (8.17) has a stationary value. The resulting Euler–Lagrange equations are known as theLagrange equations of motion: d dt/C64L /C64_/C113iÿ/C64L /C64/C113iˆ0: …8:19† These Lagrange equations can be derived from Newton’s equations of motion (that is, the second law written in di/C128erential equation form) and Newton’s equa- tions can be derived from Lagrange’s equations. Thus they are ‘equivalent.’ However, Hamilton’s principle can be applied to a wide range of physical phe- nomena, particularly those involving fields, with which Newton’s equations are not usually associated. Therefore, Hamilton’s principle is considered to be more fundamental than Newton’s equations and is often introduced as a basic postulate from which various formulations of classical dynamics are derived. Example 8.3 Electric oscillations: As an illustration of the generality of Lagrangian dynamics, we consider its application to an L/C67circuit (inductive–capacitive circuit) as shown in Fig. 8.4. At some instant of time the charge on the capacitor /C67is/C81…t†and the current flowing through the inductor is I…t†ˆ _/C81…t†. The voltage drop around the 356THE CALCULUS OF VARIATIONS Figure 8.4. L/C67circuit. circuit is, according to Kirchho/C128 ’s law LdI dt‡1 CZ I…t†dtˆ0 or in terms of /C81 L/C127/C81‡1 C/C81ˆ0: This equation is of exactly the same form as that for a simple mechanical oscil- lator: m/C127x‡kxˆ0: If the electric circuit also contains a resistor /C82, Kirchho/C128 ’s law then gives L/C127/C81‡R_/C81‡1 C/C81ˆ0; which is of exactly the same form as that for a damped oscillator m/C127x‡b_x‡kxˆ0; where bis the damping constant. By comparing the corresponding terms in these equations, an analogy between mechanical and electric quantities can be established: x displacement /C81 charge (generalized coordinate) _x velocity _/C81ˆIelectric current m mass L inductance 1=kk ˆspring constant /C67 capacitance b damping constant /C82 electric resistance 1 2m_x2kinetic energy12L_/C812energy stored in inductance 1 2mx2potential energy12/C812=Cenergy stored in capacitance If we recognize in the beginning that the charge /C81in the circuit plays the role of a generalized coordinate, and Tˆ1 2L_/C812and/C86ˆ12/C812=C, then the Langrangian L of the system is LˆTÿ/C86ˆ1 2L_/C812ÿ12/C812=C and the Lagrange equation gives L/C127/C81‡1 C/C81ˆ0; the same equation as given by Kirchho/C128 ’s law. 357HAMILTON’S PRINCIPLE Example 8.4 A bead of mass mslides freely on a frictionless wire of radius bthat rotates in a horizontal plane about a point on the circular wire with a constant angular velocity /C33. Show that the bead oscillates as a pendulum of length /C108ˆ/C103=/C332. Solution: The circular wire rotates in the xyplane about the point O, as shown in Fig. 8.5. The rotation is in the counterclockwise direction, /C67is the center of the circular wire, and the angles and /C30are as indicated. The wire rotates with an angular velocity /C33,s o/C30ˆ/C33t. Now the coordinates xandyof the bead are given by xˆbcos/C33t‡bcos…‡/C33t†; yˆbsin/C33t‡bsin…‡/C33t†; and the generalized coordinate is . The potential energy of the bead (in a hor- izontal plane) can be taken to be zero, while its kinetic energy is Tˆ1 2m…_x2‡_y2†ˆ12mb2‰/C332‡… _‡/C33†2‡2/C33…_‡/C33†cosŠ; which is also the Lagrangian of the bead. Inserting this into Lagrange’s equation d d/C64L /C64_ ÿ/C64L /C64ˆ0 we obtain, after some simplifications, /C127‡/C332sinˆ0: Comparing this equation with Lagrange’s equation for a simple pendulum of length l /C127‡…/C103=/C108†sinˆ0 358THE CALCULUS OF VARIATIONS Figure 8.5. (Fig. 8.6) we see that the bead oscillates about the line OAlike a pendulum of length /C108ˆ/C103=/C332. /C82a/C121leigh/C177/C82it/C122 method Hamilton’s principle views the motion of a dynamical system as a whole and involves a search for the path in configuration space that yields a stationary value for the action integral (8.17): IˆZt2 t1L/C113 i…t†;_/C113i…t†;t …† dtˆ0; …8:18† with /C113i…t1†ˆ/C113i…t2†ˆ0. Ordinarily it is used as a variational method to obtain Lagrange’s and Hamilton’s equations of motion, so we do not often think of it as a computational tool. But in other areas of physics variational formulationsare used in a much more active way. For example, the variational method for determining the approximate ground-state energies in quantum mechanics is very well known. We now use the Rayleigh–Ritz method to illustrate that Hamilton’s principle can be used as computational device in classical mechanics. The Rayleigh–Ritz method is a procedure for obtaining approximate solutions of problems expressed in variational form directly from the variational equation. The Lagrangian is a function of the generalized coordinates /C113s and their time derivatives _/C113s. The basic idea of the approximation method is to guess a solution for the /C113s that depends on time and a number of parameters. The parameters are then adjusted so that Hamilton’s principle is satisfied. The Rayleigh–Ritz method takes a special form for the trial solution. A complete set of functions ff i…t†gis chosen and the solution is assumed to be a linear combination of a finite number of these functions. The coecients in this linear combination are the parameters that are chosen to satisfy Hamilton’s principle (8.18). Since the variations of the /C113s 359RAYLEIGH–RIT/C90 METHOD Figure 8.6. must vanish at the endpoints of the integral, the variations of the parameter must be so chosen that this condition is satisfied. To summarize, suppose a given system can be described by the action integral IˆZt2 t1L/C113 i…t†;_/C113i…t†;t …† dt; _/C113ˆd/C113=dt: The Rayleigh–Ritz method requires the selection of a trial solution, ideally in theform /C113ˆX n iˆ1aifi…t†; …8:20† which satisfies the appropriate conditions at both the initial and final times, and where as are undetermined constant coecients and the fs are arbitrarily chosen functions. This trial solution is substituted into the action integral Iand integra- tion is performed so that we obtain an expression for the integral Iin terms of the coecients. The integral Iis then made ‘stationary’ with respect to the assumed solution by requiring that /C64I /C64aiˆ0 …8:21† after which the resulting set of nsimultaneous equations is solved for the values of the coecients ai. To illustrate this method, we apply it to two simple examples. Example 8.5A simple harmonic oscillator consists of a mass Mattached to a spring of force constant k. As a trial function we take the displacement xas a function tin the form x…t†ˆX 1 nˆ1Ansinn/C33t: For the boundary conditions we have xˆ0;tˆ0, and xˆ0;tˆ2=/C33. Then the potential energy and the kinetic energy are given by, respectively, /C86ˆ1 2kx2ˆ12kP1 nˆ1P1 mˆ1AnAmsinn/C33tsinm/C33t; Tˆ1 2M_x2ˆ12M/C332P1 nˆ1P1 mˆ1AnAmnmcosn/C33tcosm/C33t: The action Ihas the form IˆZ2=/C33 0LdtˆZ2=/C33 0…Tÿ/C86†dtˆ 2/C33X1 nˆ1…kA2 nÿMn2A2n/C332†: 360THE CALCULUS OF VARIATIONS In order to satisfy Hamilton’s principle we must choose the values of Anso as to make Ian extremum: dI dAnˆ…kÿn2/C332M†Anˆ0: The solution that meets the physics of the problem is A1ˆ0;/C332ˆk=M; or /C28ˆ2=/C33…†1=2ˆ2M=k…†1=2; Anˆ0;for nˆ2;3;etc: Example 8.6 As a second example, we consider a bead of mass Msliding freely along a wire shaped in the form of a parabola along the vertical axis and of the form yˆax2. In this case, we have LˆTÿ/C86ˆ1 2M…_x2‡_y2†ÿM/C103y ˆ12M…1‡4a2x2†_x2ÿM/C103y : We assume xˆAsin/C33t to be an approximate value for the displacement x, and then the action integral becomes IˆZ2=/C33 0LdtˆZ2=/C33 0…Tÿ/C86†dtˆA2/C332…1‡a2A2† 2ÿ/C103a() M /C33: The extremum condition, dI=dAˆ0, gives an approximate /C33: /C33ˆ2/C103ap 1‡a2A2; and the approximate period is /C28ˆ2…1‡a2A2†2/C103ap : The Rayleigh–Ritz method discussed in this section is a special case of the general Rayleigh–Ritz methods that are designed for finding approximate solu- tions of boundary-value problems by use of varitional principles, for example, the eigenvalues and eigenfunctions of the Sturm–Liouville systems. /C72amilton/C39s principle and canonical equations of motion Newton first formulated classical mechanics in the seventeenth century and it is known as Newtonian mechanics. The essential physics involved in Newtonian 361HAMILTON’S PRINCIPLE mechanics is contained in Newton’s three laws of motion, with the second law serving as the equation of motion. Classical mechanics has since been reformu-lated in a few di/C128erent forms: the Lagrange, the Hamilton, and the Hamilton– Jacobi formalisms, to name just a few. The essential physics of Lagrangian dynamics is contained in the Lagrange function Lof the dynamical system and Lagrange’s equations (the equations of motion). The Lagrangian Lis defined in terms of independent generalized coor- dinates /C113 iand the corresponding generalized velocity _/C113i. In Hamiltonian dynamics, we describe the state of a system by Hamilton’s function (or theHamiltonian) /C72defined in terms of the generalized coordinates /C113 iand the corre- sponding generalized momenta /C112i, and the equations of motion are given by Hamilton’s equations or canonical equations _/C113iˆ/C64H /C64/C112i; _/C112iˆÿ/C64H /C64/C113i; iˆ1;2;...;n: …8:22† Hamilton’s equations of motion can be derived from Hamilton’s principle. Before doing so, we have to define the generalized momentum and the Hamiltonian. The generalized momentum /C112icorresponding to /C113iis defined as /C112iˆ/C64L /C64/C113i…8:23† and the Hamiltonian of the system is defined by HˆX i/C112i_/C113iÿL: …8:24† Even though _/C113iexplicitly appears in the defining expression (8.24), /C72is a function of the generalized coordinates /C113i, the generalized momenta /C112i, and the time t, because the defining expression (8.23) can be solved explicitly for the _/C113isi n terms of /C112i;/C113i, and t. The /C113sa n d ps are now treated the same: HˆH…/C113i;/C112i;t†. Just as with the configuration space spanned by the nindependent /C113s, we can imagine a space of 2 ndimensions spanned by the 2 nvariables /C1131;/C1132;...;/C113n;/C1121;/C1122;...;/C112n. Such a space is called phase space, and is particularly useful in both statistical mechanics and the study of non-linear oscillations. Theevolution of a representative point in this space is determined by Hamilton’sequations. We are ready to deduce Hamilton’s equation from Hamilton’s principle. The original Hamilton’s principle refers to paths in configuration space, so in order to extend the principle to phase space, we must modify it such that the integrand of the action Iis a function of both the generalized coordinates and momenta and their derivatives. The action Ican then be evaluated over the paths of the system 362THE CALCULUS OF VARIATIONS point in phase space. To do this, first we solve Eq. (8.24) for L LˆX i/C112i_/C113iÿH and then substitute Linto Eq. (8.18) and we obtain IˆZt2 t1X i/C112i_/C113iÿH…/C112;/C113;t† dtˆ0; …8:25† where /C113I…t†is still varied subject to /C113i…t1†ˆ/C113i…t2†ˆ0, but /C112iis varied without such end-point restrictions. Carrying out the variation, we obtain Zt2 t1X i/C112i_/C113i‡_/C113i/C112iÿ/C64H /C64/C113i/C113iÿ/C64H /C64/C112i/C112i dtˆ0; …8:26† where the _/C113s are related to the /C113s by the relation _/C113iˆd dt/C113i: …8:27† Now we integrate the term /C112i_/C113idtby parts. Using Eq. (8.27) and the endpoint conditions on /C113i, we find that Zt2 t1X i/C112i_/C113idtˆZt2 t1X i/C112id dt/C113idt ˆZt2 t1X id dt/C112i/C113idtÿZt2 t1X i_/C112i/C113idt ˆ/C112i/C113i/C12/C12/C12/C12t2 t1ÿZt2 t1X i_/C112i/C113idt ˆÿZt2 t1X i_/C112i/C113idt: Substituting this back into Eq. (8.26), we obtain Zt2 t1X i_/C113iÿ/C64H /C64/C112i /C112iÿ _/C112i‡/C64H /C64/C113i /C113i dtˆ0: …8:28† Since we view Hamilton’s principle as a variational principle in phase space, both the/C113s and the /C112s are arbitrary, the coecients of /C113iand /C112iin Eq. (8.28) must vanish separately, which results in the 2 nHamilton’s equations (8.22). Example 8.7 Obtain Hamilton’s equations of motion for a one-dimensional harmonic oscilla- tor. 363HAMILTON’S PRINCIPLE Solution: We have Tˆ1 2m_x2; /C86ˆ12/C75x2; /C112ˆ/C64L /C64_xˆ/C64T /C64_xˆm_x; _xˆ/C112 m: Hence Hˆ/C112_xÿLˆT‡/C86ˆ1 2m/C1122‡1 2/C75x2: Hamilton’s equations _xˆ/C64H /C64/C112; _/C112ˆÿ/C64H /C64x then read _xˆ/C112 m; _/C112ˆÿ/C75x: Using the first equation, the second can be written d dt…m_x†ˆÿ /C75x orm/C127x‡/C75xˆ0 which is the familiar equation of the harmonic oscillator. /C84he modified /C72amilton/C39s principle and the /C72amilton/C177/C74acobi equation The Hamilton–Jacobi equation is the cornerstone of a general method of integrat- ing equations of motion. Before the advent of modern quantum theory, Bohr’s atomic theory was treated in terms of Hamilton–Jacobi theory. It also plays an important role in optics as well as in canonical perturbation theory. In classical mechanics books, the Hamilton–Jacobi equation is often obtained via canonical transformations. We want to show that the Hamilton–Jacobi equation can also beobtained directly from Hamilton’s principle, or, a modified Hamilton’s principle. In formulating Hamilton’s principle, we have considered the action IˆZ t2 t1L/C113 i…t†;_/C113i…t†;t …† dt; _/C113ˆd/C113=dt; taken along a path between two given positions /C113i…t1†and/C113i…t2†which the dyna- mical system occupies at given instants t1andt2. In varying the action, we com- pare the values of the action for neighboring paths with fixed ends, that is, with /C113i…t1†ˆ/C113i…t2†ˆ0. Only one of these paths corresponds to the true dynamical path for which the action has its extremum value. We now consider another aspect of the concept of action, by regarding Ias a quantity characterizing the motion along the true path, and comparing the value 364THE CALCULUS OF VARIATIONS ofIfor paths having a common beginning at /C113i…t1†, but passing through di/C128erent points at time t2. In other words we consider the action Ifor the true path as a function of the coordinates at the upper limit of integration: IˆI…/C113i;t†; where /C113iare the coordinates of the final position of the system, and tis the instant when this position is reached. If/C113i…t2†are the coordinates of the final position of the system reached at time t2, the coordinates of a point near the point /C113i…t2†can be written as /C113i…t1†‡/C113i, where /C113iis a small quantity. The action for the trajectory bringing the system to the point /C113i…t1†‡/C113idi/C128ers from the action for the trajectory bringing the system to the point /C113i…t2†by the quantity IˆZt2 t1/C64L /C64/C113i/C113i‡/C64L /C64_/C113i_/C113i dt; …8:29† where /C113iis the di/C128erence between the values of /C113itaken for both paths at the same instant t; similarly, _/C113iis the di/C128erence between the values of _/C113iat the instant t. We now integrate the second term on the right hand side of Eq. (8.25) by parts: Zt2 t1/C64L /C64_/C113i_/C113idtˆ/C64L /C64_/C113i/C113iÿZt2 t1d dt/C64L /C64_/C113i /C113idt ˆ/C112i/C113iÿZt2 t1d dt/C64L /C64_/C113i /C113idt; …8:30† where we have used the fact that the starting points of both paths coincide, hence /C113i…t1†ˆ0; the quantity /C113i…t2†is now written as just /C113i. Substituting Eq. (8.30) into Eq. (8.29), we obtain IˆX i/C112i/C113i‡Zt2 t1X i/C64L /C64/C113iÿd dt/C64L /C64_/C113i  /C113idt: …8:31† Since the true path satisfies Lagrange’s equations of motion, the integrand and,consequently, the integral itself vanish. We have thus obtained the following value for the increment of the action Idue to the change in the coordinates of the final position of the system by /C113 i(at a constant time of motion): IˆX i/C112i/C113i; …8:32† from which it follows that /C64I /C64/C113iˆ/C112i; …8:33† that is, the partial derivatives of the action with respect to the generalized co-ordinates equal the corresponding generalized momenta. 365THE MODIFIED HAMILTON’S PRINCIPLE The action Imay similarly be regarded as an explicit function of time, by considering paths starting from a given point /C113i…1†at a given instant t1, ending at a given point /C113i…2†at various times t2ˆt: IˆI…/C113i;t†: Then the total time derivative of Iis dI dtˆ/C64I /C64t‡X i/C64I /C64/C113i_/C113iˆ/C64I /C64t‡X i/C112i_/C113i: …8:34† From the definition of the action, we have dI=dtˆL. Substituting this into Eq. (8.34), we obtain /C64I /C64tˆLÿX i/C112i_/C113iˆÿH or /C64I /C64t‡H…/C113i;/C112i;t†ˆ0: …8:35† Replacing the momenta /C112iin the Hamiltonian Hby/C64I=/C64/C113ias given by Eq. (8.33), we obtain the Hamilton–Jacobi equation H…/C113i;/C64I=/C64/C113i;t†‡/C64I /C64tˆ0: …8:36† For a conservative system with stationary constraints, the time is not contained explicitly in Hamiltonian /C72,a n d Hˆ/C69(the total energy of the system). Consequently, according to Eq. (8.35), the dependence of action Ion time tis expressed by the term ÿ/C69t. Therefore, the action breaks up into two terms, one of which depends only on /C113i, and the other only on t: I…/C113i;t†ˆI/C111…/C113i†ÿ/C69t: …8:37† The function I/C111…/C113i†is sometimes called the contracted action, and the Hamilton– Jacobi equation (8.36) reduces to H…/C113i;/C64I/C111=/C64/C113i†ˆ/C69: …8:38† Example 8.8 To illustrate the method of Hamilton–Jacobi, let us consider the motion of an electron of charge ÿerevolving about an atomic nucleus of charge Ze(Fig. 8.7). As the mass Mof the nucleus is much greater than the mass mof the electron, we may consider the nucleus to remain stationary without making any very appreci- able error. This is a central force motion and so its motion lies entirely in one plane (see /C67lassical Mechanics , by Tai L. Chow, John Wiley, 1995). Employing 366THE CALCULUS OF VARIATIONS polar coordinates rand in the plane of motion to specify the position of the electron relative to the nucleus, the kinetic and potential energies are, respectively, Tˆ1 2m…_r2‡r2_2†; /C86ˆÿZe2 r: Then LˆTÿ/C86ˆ12m…_r 2‡r2_2†‡Ze2 r and /C112rˆ/C64L /C64_rˆm_r/C112 /C112ˆ/C64L /C64_ˆmr2_: The Hamiltonian /C72is Hˆ1 2m/C1122 r‡/C1122  r2/C32! ÿZe2 r: Replacing /C112rand/C112in the Hamiltonian by /C64I=/C64rand /C64I=/C64, respectively, we obtain, by Eq. (8.36), the Hamilton–Jacobi equation 1 2m/C64I /C64r2 ‡1 r2/C64I /C642"# ÿZe2 r‡/C64I /C64tˆ0: /C86ariational problems /C119ith se/C118eral independent /C118ariables The functional fin Eq. (8.1) contains only one independent variable, but very often fmay contain several independent variables. Let us now extend the theory to this case of several independent variables: IˆZZZ /C86ffu;ux;uy;uz;x;y;z†dxdydz ; …8:39† where /C86is assumed to be a bounded volume in space with prescribed values of u…x;y;z†at its boundary S;uxˆ/C64u=/C64x, and so on. Now, the variational problem 367VARIATIONAL PROBLEMS Figure 8.7. is to find the function u…x;y;z†for which Iis stationary with respect to small changes in the functional form u…x;y;z†. Generalizing Eq. (8.2), we now let u…x;y;z;/C34†ˆu…x;y;z;0†‡/C34/C17…x;y;z†; …8:40† where /C17…x;y;z†is an arbitrary well-behaved (that is, di/C128erentiable) function which vanishes at the boundary S. Then we have, from Eq. (8.40), ux…x;y;z;/C34†ˆux…x;y;z;0†‡/C34/C17x; and similar expressions for uy;uz;a n d /C64I /C64/C34/C12/C12/C12/C12 /C34ˆ0ˆZZZ /C86/C64f /C64u/C17‡/C64f /C64ux/C17x‡/C64f /C64uy/C17y‡/C64f /C64uz/C17z dxdydz ˆ0: We next integrate each of the terms …/C64f=/C64ui†/C17iusing ‘integration by parts’ and the integrated terms vanish at the boundary as required. After some simplifications, we finally obtain ZZZ /C86/C64f /C64uÿ/C64 /C64x/C64f /C64uxÿ/C64 /C64y/C64f /C64uyÿ/C64 /C64z/C64f /C64uz/C26/C27 /C17…x;y;z†dxdydz ˆ0: Again, since /C17…x;y;z†is arbitrary, the term in the braces may be set equal to zero, and we obtain the Euler–Lagrange equation: /C64f /C64uÿ/C64 /C64x/C64f /C64uxÿ/C64 /C64y/C64f /C64uyÿ/C64 /C64z/C64f /C64uzˆ0: …8:41† Note that in Eq. (8.41) /C64=/C64xis a partial derivative, in that yandzare constant. But/C64=/C64xis also a total derivative in that it acts on implicit xdependence and on explicit xdependence: /C64 /C64x/C64f /C64uxˆ/C642f /C64x/C64ux‡/C642f /C64u/C64uxux‡/C642f /C64u2x‡/C642f /C64uy/C64uxuxy‡/C642f /C64uz/C64uxuxz: …8:42† Example 8.9 The Schro /C200dinger wave equation. The equations of motion of classical mechanics are the Euler–Lagrange di/C128erential equations of Hamilton’s principle. Similarly,the Schro /C200dinger equation, the basic equation of quantum mechanics, is also a Euler–Lagrange di/C128erential equation of a variational principle the form of which is, in the case of a system of /C78particles, the following Z Ld/C28ˆ0; …8:43† 368THE CALCULUS OF VARIATIONS with LˆX/C78 iˆ1p2 2mi/C64/C32/C42 /C64xi/C64/C32 /C64xi‡/C64/C32/C42 /C64yi/C64/C32 /C64yi‡/C64/C32/C42 /C64zi/C64/C32 /C64zi ‡/C86/C32/C42/C32 …8:44† and the constraint Z /C32/C42/C32d/C28ˆ1; …8:45† where miis the mass of particle I,/C86is the potential energy of the system, and d/C28is a volume element of the 3 /C78-dimensional space. Condition (8.45) can be taken into consideration by introducing a Lagrangian multiplier ÿ/C69: Z …Lÿ/C69/C32/C42/C32†d/C28ˆ0: …8:46† Performing the variation we obtain the Schro /C200dinger equation for a system of /C78 particles X/C78 iˆ1p2 2mi/C1142 i/C32‡…/C69ÿ/C86†/C32ˆ0; …8:47† where /C1142 iis the Laplace operator relating to particle i. Can you see that Eis the energy parameter of the system/C63 If we use the Hamiltonian operator ^H, Eq. (8.47) can be written as ^H/C32ˆ/C69/C32: …8:48† From this we obtain for E /C69ˆZ /C32/C42H/C32d/C28 Z /C32/C42/C32d/C28: …8:49† Through partial integration we obtain Z Ld/C28ˆZ /C32/C42H/C32d/C28 and thus the variational principle can be formulated in another way: R /C32/C42…Hÿ/C69†/C32d/C28ˆ0. Problems 8.1 As a simple practice of using varied paths and the extremum condition, we consider the simple function y…x†ˆxand the neighboring paths 369PROBLEMS y…/C34;x†ˆx‡/C34sinx. Draw these paths in the xyplane between the limits xˆ0a n d xˆ2for/C34ˆ0 for two di/C128erent non-vanishing values of /C34. If the integral I…/C34†is given by I…/C34†ˆZ2 0…dy=dx†2dx; show that the value of I…/C34†is always greater than I…0†, no matter what value of/C34(positive or negative) is chosen. This is just condition (8.4). 8.2 ( a) Show that the Euler–Lagrange equation can be written in the form d dxfÿy0/C64f /C64y0 ÿ/C64f /C64xˆ0: This is often called the second form of the Euler–Lagrange equation. (b)I ffdoes not involve xexplicitly, show that the Euler–Lagrange equation can be integrated to yield fÿy0/C64f /C64y0ˆc; where cis an integration constant. 8.3 As shown in Fig. 8.8, a curve /C67joining points …x1;y1†and…x2;y2†is revolved about the x-axis. Find the shape of the curve such that the surface thus generated is a minimum. 8.4 A geodesic is a line that represents the shortest distance between two points. Find the geodesic on the surface of a sphere. 8.5 Show that the geodesic on the surface of a right circular cylinder is a helix. 8.6 Find the shape of a heavy chain which minimizes the potential energy while the length of the chain is constant. 8.7 A wedge of mass Mand angle slides freely on a horizontal plane. A particle of mass mmoves freely on the wedge. Determine the motion of the particle as well as that of the wedge (Fig. 8.9). 370THE CALCULUS OF VARIATIONS Figure 8.8. 8.8 Use the Rayleigh–Ritz method to analyze the forced oscillations of a har- monic oscillation: m/C127x‡kxˆ/C700sin/C33t: 8.9 A particle of mass mis attracted to a fixed point Oby an inverse square force /C70rˆÿk=r2(Fig. 8.10). Find the canonical equations of motion. 8.10 Set up the Hamilton–Jacobi equation for the simple harmonic oscillator. 371PROBLEMS Figure 8.9. Figure 8.10. 9 The Laplace transformation The Laplace transformation method is generally useful for obtaining solutions of linear di/C128erential equations (both ordinary and partial). It enables us to reduce a di/C128erential equation to an algebraic equation, thus avoiding going to the trouble of finding the general solution and then evaluating the arbitrary constants. This procedure or technique can be extended to systems of equations and to integral equations, and it often yields results more readily than other techniques. In this chapter we shall first define the Laplace transformation, then evaluate the trans- formation for some elementary functions, and finally apply it to solve some simple physical problems. /C68efinition of the Lapace transform The Laplace transform L‰f…x†Šof a function f…x†is defined by the integral L‰f…x†Š ˆZ1 0eÿ/C112xf…x†dxˆ/C70…/C112†; …9:1† whenever this integral exists. The integral in Eq. (9.1) is a function of the para- meter pand we denote it by /C70…/C112†. The function /C70…/C112†is called the Laplace trans- form of f…x†. We may also look upon Eq. (9.1) as a definition of a Laplace transform operator Lwhich tranforms f…x†in to /C70…/C112†. The operator Lis linear, since from Eq. (9.1) we have L‰c1f…x†‡c2/C103…x†Š ˆZ1 0eÿ/C112xfc1f…x†‡c2/C103…x†gdx ˆc1Z1 0eÿ/C112xf…x†dx‡c2Z1 0eÿ/C112x/C103…x†dx ˆc1L‰f…x†Š ‡c2L‰/C103…x†Š; 372 where c1andc2are arbitrary constants and /C103…x†is an arbitrary function defined forx/C620. The inverse Laplace transform of /C70…/C112†is a function f…x†such that L‰f…x†Š ˆ/C70…/C112†. We denote the operation of taking an inverse Laplace transform byLÿ1: Lÿ1‰/C70…/C112†Š ˆf…x†: …9:2† That is, we operate algebraically with the operators LandLÿ1, bringing them from one side of an equation to the other side just as we would in writing axˆb implies xˆaÿ1b. To illustrate the calculation of a Laplace transform, let us consider the following simple example. Example 9.1 Find L‰eaxŠ, where ais a constant. Solution: The transform is L‰eaxŠˆZ1 0eÿ/C112xeaxdxˆZ1 0eÿ…/C112ÿa†xdx: For/C112a, the exponent on eis positive or zero and the integral diverges. For /C112/C62a, the integral converges: L‰eaxŠˆZ1 0eÿ/C112xeaxdxˆZ1 0eÿ…/C112ÿa†xdxˆeÿ…/C112ÿa†x ÿ…/C112ÿa†/C12/C12/C12/C121 0ˆ1 /C112ÿa: This example enables us to investigate the existence of Eq. (9.1) for a general function f…x†. /C69/C120istence of Laplace transforms We can prove that: (1) if f…x†is piecewise continuous on every finite interval 0 xX, and (2) if we can find constants Mand asuch that jf…x†j MeaxforxX, then L‰f…x†Šexists for /C112/C62a. A function f…x†which satisfies condition (2) is said to be of exponential order asx!1 ; this is mathematician’s jargon/C33 These are sucient conditions on f…x†under which we can guarantee the existence of L‰f…x†Š. Under these conditions the integral converges for /C112/C62a: ZX 0f…x†eÿ/C112xdx/C12/C12/C12/C12/C12/C12/C12/C12Z X 0f…x†jj eÿ/C112xdxZX 0Meaxeÿ/C112xdx MZ1 0eÿ…/C112ÿa†xdxˆM /C112ÿa: 373E/C88ISTENCE OF LAPLACE TRANSFORMS This establishes not only the convergence but the absolute convergence of the integral defining L‰f…x†Š. Note that M=…/C112ÿa†tends to zero as /C112!1 . This shows that lim /C112!1/C70…/C112†ˆ0 …9:3† for all functions /C70…/C112†ˆL‰f…x†Šsuch that f…x†satisfies the foregoing conditions (1) and (2). It follows that if lim /C112!1/C70…/C112†6 ˆ0,/C70…/C112†cannot be the Laplace trans- form of any function f…x†. It is obvious that functions of exponential order play a dominant role in the use of Laplace transforms. One simple way of determining whether or not a specifiedfunction is of exponential order is the following one: if a constant bexists such that lim x!1eÿbxf…x†jj/C104/C105 …9:4† exists, the function f…x†is of exponential order (of the order of eÿbx†. To see this, let the value of the above limit be /C756ˆ0. Then, when xis large enough, jeÿbxf…x†j can be made as close to /C75as possible, so certainly jeÿbxf…x†j<2/C75: Thus, for suciently large x, jf…x†j<2/C75ebx or jf…x†j<Mebx;with Mˆ2/C75: On the other hand, if lim x!1‰eÿcxf…x†jj Š ˆ1 … 9:5† for every fixed c, the function f…x†is not of exponential order. To see this, let us assume that bexists such that jf…x†j<Mebxfor xX from which it follows that jeÿ2bxf…x†j<Meÿbx: Then the choice of cˆ2bwould give us jeÿcxf…x†j<Meÿbx, and eÿcxf…x†!0a s x!1 which contradicts Eq. (9.5). Example 9.2 Show that x3is of exponential order as x!1 . 374THE LAPLACE TRANSFORMATION Solution: We have to check whether or not lim x!1eÿbxx3 ˆlim x!1x3 ebx exists. Now if b/C620, then L’Hospital’s rule gives lim x!1eÿbxx3 ˆlim x!1x3 ebxˆlim x!13x2 bebxˆlim x!16x b2ebxˆlim x!16 b3ebxˆ0: Therefore x3is of exponential order as x!1 . Laplace transforms of some elementar/C121 functions Using the definition (9.1) we now obtain the transforms of polynomials, expo- nential and trigonometric functions. (1)f…x†ˆ1 for x/C620. By definition, we have L‰1ŠˆZ1 0eÿ/C112xdxˆ1 /C112; /C112/C620: (2)f…x†ˆxn, where nis a positive integer. By definition, we have L‰xnŠˆZ1 0eÿ/C112xxndx: Using integration by parts: Z u/C1180dxˆu/C118ÿZ /C118u0dx with uˆxn;d/C118ˆ/C1180dxˆeÿ/C112xdxˆÿ … 1=/C112†d…eÿ/C112x†;/C118ˆÿ … 1=/C112†eÿ/C112x; we obtain Z1 0eÿ/C112xxndxˆÿxneÿ/C112x /C112 1 0‡n/C112Z 1 0eÿ/C112xxnÿ1dx: For/C112/C620 and n/C620, the first term on the right hand side of the above equation is zero, and so we have Z1 0eÿ/C112xxndxˆn/C112Z 1 0eÿ/C112xxnÿ1dx 375LAPLACE TRANSFORMS OF ELEMENTARY FUNCTIONS or L‰xnŠˆn /C112L‰xnÿ1Š from which we may obtain for n/C621 L‰xnÿ1Šˆnÿ1 /C112L‰xnÿ2Š: Iteration of this process yields L‰xnŠˆn…nÿ1†…nÿ2† 21 /C112nL‰x0Š: By (1) above we have L‰x0ŠˆL‰1Šˆ1=/C112: Hence we finally have L‰xnŠˆn/C33 /C112n‡1; /C112/C620: (3)f…x†ˆeax, where ais a real constant. L‰eaxŠˆZ1 0eÿ/C112xeaxdxˆ1 /C112ÿa; where /C112/C62afor convegence. (For details, see Example 9.1.) (4)f…x†ˆsinax, where ais a real constant. L‰sinaxŠˆZ1 0eÿ/C112xsinaxdx : UsingZ u/C1180dxˆu/C118ÿZ /C118u0dx with uˆeÿ/C112x;d/C118ˆÿd…cosax†=a; and Z emxsinnxdx ˆemx…msinnxÿncosnx† n2‡m2 (you can obtain this simply by using integration by parts twice) we obtain L‰sinaxŠˆZ1 0eÿ/C112xsinaxdx ˆeÿ/C112x…ÿ/C112sinaxÿacosax /C1122‡a2 1 0: Since pis positive, eÿ/C112x!0a s x!1 , but sin axand cos axare bounded as x!1 , so we obtain L‰sinaxŠˆ0ÿ1…0ÿa† /C1122‡a2ˆa /C1122‡a2;/C112/C620: 376THE LAPLACE TRANSFORMATION (5)f…x†ˆcosax, where ais a real constant. Using the result Z emxcosnxdx ˆemx…mcosnx‡nsinmx† n2‡m2; we obtain L‰cosaxŠˆZ1 0eÿ/C112xcosaxdx ˆ/C112 /C1122‡a2; /C112/C620: (6)f…x†ˆsinhax, where ais a real constant. Using the linearity property of the Laplace transform operator L, we obtain L‰cosh axŠˆLeax‡eÿax 2 ˆ1 2L‰eaxЇ12L‰e ÿaxŠ ˆ12 1 /C112ÿa‡1 /C112‡a ˆ/C112 /C1122ÿa2: (7)f…x†ˆxk, where k/C62ÿ1. By definition we have L‰xkŠˆZ1 0eÿ/C112xxkdx: Let/C112xˆu, then dxˆ/C112ÿ1du;xkˆuk=/C112k, and so L‰xkŠˆZ1 0eÿ/C112xxkdxˆ1 /C112k‡1Z1 0ukeÿuduˆÿ…k‡1† /C112k‡1: Note that the integral defining the gamma function converges if and only if k/C62ÿ1. The following example illustrates the calculation of inverse Laplace transforms which is equally important in solving di/C128erential equations. Example 9.3 Find …a†Lÿ15 /C112‡2 ; …b†Lÿ11 /C112s ;s/C620: Solution: …a†Lÿ15 /C112‡2 ˆ5Lÿ11 /C112‡2 : 377LAPLACE TRANSFORMS OF ELEMENTARY FUNCTIONS Recall L‰eaxŠˆ1=…/C112ÿa†, hence Lÿ1‰1=…/C112ÿa†Š ˆeax. It follows that Lÿ15 /C112‡2 ˆ5Lÿ11 /C112‡2 ˆ5eÿ2x: (b) Recall L‰xkŠˆZ1 0eÿ/C112xxkdxˆ1 /C112k‡1Z1 0ukeÿuduˆÿ…k‡1† /C112k‡1: From this we have Lxk ÿ…k‡1†"# ˆ1 /C112k‡1; hence Lÿ11 /C112k‡1 ˆxk ÿ…k‡1†: If we now let k‡1ˆs, then Lÿ11 /C112s ˆxsÿ1 ÿ…s†: /C83hifting (or translation) theorems In practical applications, we often meet functions multiplied by exponential fac- tors. If we know the Laplace transform of a function, then multiplying it by an exponential factor does not require a new computation as shown by the following theorem. /C84he /C174rst shifting theorem IfL‰f…x†Šˆ/C70…/C112†;/C112/C62b;t/C104en L ‰eatf…x†Š ˆ/C70…/C112ÿa†;/C112/C62a‡b. Note that /C70…/C112ÿa†denotes the function /C70…/C112†‘shifted’ a units to the right. Hence the theorem is called the shifting theorem. The proof is simple and straightforward. By definition (9.1) we have L‰f…x†Š ˆZ1 0eÿ/C112xf…x†dxˆ/C70…/C112†: Then Leaxf…x† ‰Š ˆZ1 0eÿ/C112xfeaxf…x†gdxˆZ1 0eÿ…/C112ÿa†xf…x†dxˆ/C70…/C112ÿa†: The following examples illustrate the use of this theorem. 378THE LAPLACE TRANSFORMATION Example 9.4 Show that: …a†Leÿaxxn‰Š ˆn/C33 …/C112‡a†n‡1; /C112/C62ÿa; …b†Leÿaxsinbx ‰Š ˆb …/C112‡a†2‡b2; /C112/C62ÿa: Solution: (a) Recall L‰xnŠˆn/C33=/C112n‡1; /C112/C620; the shifting theorem then gives L‰eÿaxxnŠˆn/C33 …/C112‡a†n‡1; /C112/C62ÿa: (b) Since L‰sinaxŠˆa /C1122‡a2; it follows from the shifting theorem that L‰eÿaxsinbxŠˆb …/C112‡a†2‡b2; /C112/C62ÿa: Because of the relationship between Laplace transforms and inverse Laplace transforms, any theorem involving Laplace transforms will have a corresponding theorem involving inverse Lapace transforms. Thus If Lÿ1/C70…/C112†‰Š ˆ f…x†;t/C104en Lÿ1/C70…/C112ÿa† ‰Š ˆ eaxf…x†: /C84he second shifting theorem This second shifting theorem involves the shifting xvariable and states that /C71iven L‰f…x†Š ˆ/C70…/C112†,where f…x†ˆ0forx<0;and if /C103…x†ˆf…xÿa†, then L‰/C103…x†Š ˆeÿa/C112L‰f…x†Š: To prove this theorem, let us start with /C70…/C112†ˆLf…x†‰Š ˆZ1 0eÿ/C112xf…x†dx from which it follows that eÿa/C112/C70…/C112†ˆeÿa/C112Lf…x†‰Š ˆZ1 0eÿ/C112…x‡a†f…x†dx: 379SHIFTING (OR TRANSLATION) THEOREMS Letuˆx‡a, then eÿa/C112/C70…/C112†ˆZ1 0eÿ/C112…x‡a†f…x†dxˆZ1 0eÿ/C112uf…uÿa†du ˆZa 0eÿ/C112u0du‡Z1 aeÿ/C112uf…uÿa†du ˆZ1 0eÿ/C112u/C103…u†duˆL/C103…u†‰Š : Example 9.5 Show that given f…x†ˆxfor x0 0f o r x<0;/C26 and if /C103…x†ˆ0; for x<5 xÿ5;for x5/C26 then L‰/C103…x†Š ˆeÿ5/C112=/C1122: Solution: We first notice that /C103…x†ˆf…xÿ5†: Then the second shifting theorem gives L‰/C103…x†Š ˆeÿ5/C112L‰xŠˆeÿ5/C112=/C1122: /C84he unit step function It is often possible to express various discontinuous functions in terms of the unitstep function, which is defined as U…xÿa†ˆ0x<a 1xa:/C26 Sometimes it is convenient to state the second shifting theorem in terms of the unit step function: Iff…x†ˆ0forx<0and L‰f…x†Š ˆ/C70…/C112†,then L‰U…xÿa†f…xÿa†Š ˆe ÿa/C112/C70…/C112†: 380THE LAPLACE TRANSFORMATION The proof is straightforward: LU…xÿa†f…xÿa† ‰Š ˆZ1 0eÿ/C112xU…xÿa†f…xÿ†dx ˆZa 0eÿ/C112x0dx‡Z1 aeÿ/C112xf…xÿa†dx: Letxÿaˆu, then LU…xÿa†f…xÿa† ‰Š ˆZ1 aeÿ/C112xf…xÿa†dx ˆZ1 aeÿ/C112…u‡a†f…u†duˆeÿa/C112Z1 aeÿ/C112uf…u†duˆeÿa/C112/C70…/C112†: The corresponding theorem involving inverse Laplace transforms can be stated as Iff…x†ˆ0forx<0and Lÿ1‰/C70…/C112†Š ˆf…x†/C44 then Lÿ1‰eÿa/C112/C70…/C112†Š ˆU…xÿa†f…xÿa†: Laplace transform of a periodic function Iff…x†is a periodic function of period P/C620, that is, if f…x‡P†ˆf…x†, then L‰f…x†Š ˆ1 1ÿeÿ/C112PZP 0eÿ/C112xf…x†dx: To prove this, we assume that the Laplace transform of f…x†exists: L‰f…x†Š ˆZ1 0eÿ/C112xf…x†dxˆZP 0eÿ/C112xf…x†dx‡Z2P Peÿ/C112xf…x†dx ‡Z3P 2Peÿ/C112xf…x†dx‡ : On the right hand side, let xˆu‡Pin the second integral, xˆu‡2Pin the third integral, and so on, we then have Lf…x†‰Š ˆZP 0eÿ/C112xf…x†dx‡ZP 0eÿ/C112…u‡P†f…u‡P†du ‡ZP 0eÿ/C112…u‡2P†f…u‡2P†du‡ : 381LAPLACE TRANSFORM OF A PERIODIC FUNCTION But f…u‡P†ˆf…u†;f…u‡2P†ˆf…u†;etc:Also, let us replace the dummy variable ubyx, then the above equation becomes Lf…x†‰Š ˆZP 0eÿ/C112xf…x†dx‡ZP 0eÿ/C112…x‡P†f…x†dx‡ZP 0eÿ/C112…x‡2P†f…x†dx‡ ˆZP 0eÿ/C112xf…x†dx‡eÿ/C112PZP 0eÿ/C112xf…x†dx‡eÿ2/C112PZP 0eÿ/C112xf…x†dx‡ ˆ…1‡eÿ/C112P‡eÿ2/C112P‡ †ZP 0eÿ/C112xf…x†dx ˆ1 1ÿeÿ/C112PZP 0eÿ/C112xf…x†dx: Laplace transforms of deri/C118ati/C118es Iff…x†is a continuous for x0, and f0…x†is piecewise continuous in every finite interval 0 xk, and if jf…x†j Mebx(that is, f…x†is of exponential order), then L‰f0…x†Š ˆ/C112L f…x†‰Š ÿ f…0†;/C112/C62b: We may employ integration by parts to prove this result:Z ud/C118ˆu/C118ÿZ /C118du with uˆeÿ/C112x;and d/C118ˆf0…x†dx; L‰f0…x†Š ˆZ1 0eÿ/C112xf0…x†dxˆ‰eÿ/C112xf…x†Š1 0ÿZ1 0…ÿ/C112†eÿ/C112xf…x†dx: Since jf…x†j Mebxfor suciently large x, then jf…x†eÿ/C112xjMe…bÿ/C112†for su- ciently large x.I f /C112/C62b, then Me…bÿ/C112†!0a s x!1 ; and eÿ/C112xf…x†!0a s x!1 . Next, f…x†is continuous at xˆ0, and so eÿ/C112xf…x†!f…0†asx!0. Thus, the desired result follows: L‰f0…x†Š ˆ/C112L‰f…x†Š ÿf…0†;/C112/C62b: This result can be extended as follows: Iff…x†is such that f…nÿ1†…x†is continuous and f…n†…x†piecewise continuous in every interval 0 xkand furthermore, if f…x†;f0…x†;...;f…n†…x†are of exponential order for 0 /C62k, then L‰f…n†…x†Š ˆ/C112nL‰f…x†Š ÿ/C112nÿ1f…0†ÿ/C112nÿ2f0…0†ÿÿ f…nÿ1†…0†: Example 9.6 Solve the initial value problem: y00‡yˆ0;y…0†ˆy0…0†ˆ0;and f…t†ˆ0f o r t<0 but f…t†ˆ1f o r t0: 382THE LAPLACE TRANSFORMATION Solution: Note that y0ˆdy=dt. We know how to solve this simple di/C128erential equation, but as an illustration we now solve it using Laplace transforms. Taking both sides of the equation we obtain L‰y00ЇL‰yŠˆL‰1Š;…L‰fŠˆL‰1І: Now L‰y00Šˆ/C112L‰y0Šÿy0…0†ˆ/C112f/C112L‰yŠÿy…0†g ÿ y0…0† ˆ/C1122L‰yŠÿ/C112y…0†ÿy0…0† ˆ/C1122L‰yŠ and L‰1Šˆ1=/C112: The transformed equation then becomes /C1122Ly‰Š ‡ Ly‰Š ˆ 1=/C112 or Ly‰Š ˆ1 /C112…/C1122‡1†ˆ1 /C112ÿ/C112 /C1122‡1; therefore yˆLÿ11 /C112 ÿLÿ1 /C112 /C1122‡1 : We find from Eqs. (9.6) and (9.10) that Lÿ11 /C112 ˆ1 and Lÿ1/C112 /C1122‡1 ˆcost: Thus, the solution of the initial problem is yˆ1ÿcostfor t0; yˆ0 for t<0: Laplace transforms of functions defined b/C121 integrals If/C103…x†ˆRx 0f…u†du,and if Lf…x†‰Š ˆ /C70…/C112†,thenL/C103…x†‰Š ˆ /C70…/C112†=/C112. Similarly, if Lÿ1/C70…/C112†‰Š ˆ f…x†, then Lÿ1/C70…/C112†=/C112 ‰Š ˆ /C103…x†: It is easy to prove this. If /C103…x†ˆRx 0f…u†du, then /C103…0†ˆ0;/C1030…x†ˆf…x†. Taking Laplace transform, we obtain L‰/C1030…x†Š ˆL‰f…x†Š 383FUNCTIONS DEFINED BY INTEGRALS but L‰/C1030…x†Š ˆ/C112L‰/C103…x†Š ÿ/C103…0†ˆ/C112L‰/C103…x†Š and so /C112L‰/C103…x†Š ˆL‰f…x†Š;orL‰/C103…x†Š ˆ1 /C112Lf…x†‰Š ˆ/C70…/C112† /C112: From this we have Lÿ1/C70…/C112†=/C112 ‰Š ˆ /C103…x†: Example 9.7 If/C103…x†ˆRu 0sinau du , then L‰/C103…x†Š ˆLZu 0sinau du ˆ1 /C112Lsinau‰Š ˆa /C112…/C1122‡a2†: /C65 note on integral transformations A Laplace transform is one of the integral transformations. The integral trans- formation T‰f…x†Šof a function f…x†is defined by the integral equation Tf…x†‰Š ˆZb af…x†/C75…/C112;x†dxˆ/C70…/C112†; …9:6† where /C75…/C112;x†, a known function of pand x, is called the kernel of the transfor- mation. In the application of integral transformations to the solution of bound- ary-value problems, we have so far made use of five di/C128erent kernels: Laplace transform: /C75…/C112;x†ˆeÿ/C112x,a n d aˆ0;bˆ1 : L‰f…x†Š ˆZ1 0eÿ/C112xf…x†dxˆ…/C112†: Fourier sine and cosine transforms: /C75…/C112;x†ˆsin/C112xor cos px, and aˆ0;bˆ1 : /C70‰f…x†Š ˆZ1 0f…x†/C26sin…/C112x† cos…/C112x†dxˆ/C70…/C112†: Complex Fourier transform: /C75…/C112;x†ˆei/C112x;andaˆÿ 1 ,bˆÿ 1 : /C70‰f…x†Š ˆZ1 ÿ1ei/C112xf…x†dxˆ/C70…/C112†: 384THE LAPLACE TRANSFORMATION Hankel transform: /C75…/C112;x†ˆx/C74n…/C112x†;aˆ0;bˆ1 , where /C74n…/C112x†is the Bessel function of the first kind of order n: H‰f…x†Š ˆZ1 0f…x†x/C74n…x†dxˆ/C70…/C112†: Mellin transform: /C75…/C112;x†ˆx/C112ÿ1,a n d aˆ0;bˆ1 : M‰f…x†Š ˆZ1 0f…x†x/C112ÿ1dxˆ/C70…/C112†: The Laplace transform has been the subject of this chapter, and the Fouier transform was treated in Chapter 4. It is beyond the scope of this book to include Hankel and Mellin transformations. Problems 9.1 Show that: (a)et2is not of exponential order as x!1 . (b) sin et2is of exponential order as x!1 . 9.2 Show that: (a)L‰sinhaxŠˆa /C1122ÿa2; /C112/C620: (b)L‰3x4ÿ2x3=2‡6Šˆ72 /C1125ÿ3p 2/C1125=2‡6 /C112. (c)L‰sinxcosxŠˆ1=…/C1122‡4†: (d)I f f…x†ˆx;0<x<4 5;x/C624;/C26 then L‰f…x†ˆ1 /C1122‡eÿ4/C112 /C112ÿeÿ4/C112 /C1122: 9.4 Show that L‰U…xÿa†Š ˆeÿa/C112=/C112;/C112/C620: 9.5 Find the Laplace transform of H…x†, where H…x†ˆx; 5;0<x<4 x/C624:/C26 9.5 Let f…x†be the rectified sine wave of period Pˆ2: f…x†ˆsinx;0<x< 0; x<2:/C26 Find the Laplace transform of f…x†. 385PROBLEMS 9.6 Find Lÿ1 15 /C1122‡4/C112‡13 : 9.7 Prove that if f0…x†is continuous and f00…x†is piecewise continuous in every finite interval 0 xkand if f…x†andf0…x†are of exponential order for x/C62k, then L‰fF…x†Š ˆ/C1122L‰f…x†Š ÿ/C112f…0†ÿf0…0†: (Hint: Use (9.19) with f0…x†in place of f…x†andf00…x†in place of f0…x†.) 9.8 Solve the initial problem y00…t†‡/C122y…t†ˆAsin/C33t;y…0†ˆ1;y0…0†ˆ0. 9.9 Solve the initial problem yF…t†ÿy0…t†ˆsintsubject to y…0†ˆ2; y…0†ˆ0; y00…0†ˆ1: 9.10 Solve the linear simultaneous di/C128erential equation with constant coecients y00‡2yÿxˆ0; x00‡2xÿyˆ0; subject to x…0†ˆ2;y…0†ˆ0, and x0…0†ˆy0…0†ˆ0, where xand yare the dependent variables and tis the independent variable. 9.11 Find LZ1 0cosau du : 9.12. Prove that if L‰f…x†Š ˆ/C70…/C112†then L‰f…ax†Š ˆ1 a/C70/C112a : Similarly if L ÿ1‰/C70…/C112†Š ˆf…x†then Lÿ1/C70/C112a/C104/C105 ˆaf…ax†: 386THE LAPLACE TRANSFORMATION 10 Partial di/C128erential e/C113uations We have met some partial di/C128erential equations in previous chapters. In this chapter we will study some elementary methods of solving partial di/C128erential equations which occur frequently in physics and in engineering. In general, the solution of partial di/C128erential equations presents a much more dicult problem than the solution of ordinary di/C128erential equations. A complete discussion of the general theory of partial di/C128erential equations is well beyond the scope of this book. We therefore limit ourselves to a few solvable partial di/C128erential equations that are of physical interest. Any equation that contains an unknown function of two or more variables and its partial derivatives with respect to these variables is called a partial di/C128erentialequation, the order of the equation being equal to the order of the highest partial derivatives present. For example, the equations 3y 2/C64u /C64x‡/C64u /C64yˆ2u;/C642u /C64x/C64yˆ2xÿy are typical partial di/C128erential equations of the first and second orders, respec- tively, xandybeing independent variables and u…x;y†the function to be found. These two equations are linear, because both uand its derivatives occur only to the first order and products of uand its derivatives are absent. We shall not consider non-linear partial di/C128erential equations. We have seen that the general solution of an ordinary di/C128erential equation contains arbitrary constants equal in number to the order of the equation. But the general solution of a partial di/C128erential equation contains arbitrary functions (equal in number to the order of the equation). After the particular choice of the arbitrary functions is made, the general solution becomes a particular solution. The problem of finding the solution of a given di/C128erential equation subject to given initial conditions is called a boundary-value problem or an initial-value 387 problem. We have seen already that such problems often lead to eigenvalue problems. Linear second-order partial di/C128erential equations Many physical processes can be described to some degree of accuracy by linear second-order partial di/C128erential equations. For simplicity, we shall restrict our discussion to the second-order linear partial di/C128erential equation in two indepen- dent variables, which has the general form A/C642u /C64x2‡B/C642u /C64x/C64y‡C/C642u /C64y2‡D/C64u /C64x‡/C69/C64u /C64y‡/C70uˆ/C71; …10:1† where A;B;C;...;/C71may be dependent on variables xandy. If/C71is a zero function, then Eq. (10.1) is called homogeneous; otherwise it is said to be non-homogeneous. If u1;u2;...;unare solutions of a linear homoge- neous partial di/C128erential equation, then c1u1‡c2u2‡‡ cnunis also a solution, where c1;c2;...are constants. This is known as the superposition principle; it does not apply to non-linear equations. The general solution of a linear non-homo-geneous partial di/C128erential equation is obtained by adding a particular solution of the non-homogeneous equation to the general solution of the homogeneous equation. The homogeneous form of Eq. (10.1) resembles the equation of a general conic: ax 2‡bxy‡cy2‡dx‡ey‡fˆ0: We thus say that Eq. (10.1) is of elliptic hyperbolic parabolic9 >= >;type whenB2ÿ4AC<0 B2ÿ4AC/C620 B2ÿ4ACˆ08 >< >:: For example, according to this classification the two-dimensional Laplace equation /C642u /C64x2‡/C642u /C64y2ˆ0 is of elliptic type ( AˆCˆ1;BˆDˆ/C69ˆ/C70ˆ/C71ˆ0†, and the equation /C642u /C64x2ÿ 2/C642u /C64y2ˆ0… is a real constant † is of hyperbolic type. Similarly, the equation /C642u /C64x2ÿ /C64u /C64yˆ0… is a real constant † is of parabolic type. 388PARTIAL DIFFERENTIAL EQUATIONS We now list some important linear second-order partial di/C128erential equations that are of physical interest and we have seen already: (1) Laplace’s equation: /C1142uˆ0; …10:2† where /C1142is the Laplacian operator. The function umay be the electrostatic potential in a charge-free region. It may be the gravitational potential in a region containing no matter or the velocity potential for an incompressible fluid with no sources or sinks. (2) Poisson’s equation: /C1142uˆ/C26…x;y;z†; …10:3† where the function /C26…x;y;z†is called the source density. For example, if urepre- sents the electrostatic potential in a region containing charges, then /C26is propor- tional to the electrical charge density. Similarly, for the gravitational potential case, /C26is proportional to the mass density in the region. (3) Wave equation: /C1142uˆ1 /C1182/C642u /C64t2; …10:4† transverse vibrations of a string, longitudinal vibrations of a beam, or propaga-tion of an electromagnetic wave all obey this same type of equation. For a vibrat- ing string, urepresents the displacement from equilibrium of the string; for a vibrating beam, uis the longitudinal displacement from the equilibrium. Similarly, for an electromagnetic wave, umay be a component of electric field /C69or magnetic field /C72. (4) Heat conduction equation: /C64u /C64tˆ /C1142u; …10:5† where uis the temperature in a solid at time t. The constant is called the di/C128usivity and is related to the thermal conductivity, the specific heat capacity, and the mass density of the object. Eq. (10.5) can also be used as a di/C128usion equation: uis then the concentration of a di/C128using substance. It is obvious that Eqs. (10.2)–(10.5) all are homogeneous linear equations with constant coecients. Example 10.1 Laplace’s equation: arises in almost all branches of analysis. A simple example can be found from the motion of an incompressible fluid. Its velocity /C118…x;y;z;t† and the fluid density /C26…x;y;z;t†must satisfy the equation of continuity: /C64/C26 /C64t‡/C114… /C26/C118†ˆ0: 389LINEAR SECOND-ORDER PDEs If/C26is constant we then have /C114/C118ˆ0: If, furthermore, the motion is irrotational, the velocity vector can be expressed as the gradient of a scalar function /C86: /C118ˆÿ /C114 /C86; and the equation of continuity becomes Laplace’s equation: /C114/C118ˆ /C114  …ÿ/C114 /C86†ˆ0;or/C1142/C86ˆ0: The scalar function /C86is called the velocity potential. Example 10.2 Poisson’s equation: The electrostatic field provides a good example of Poisson’s equation. The electric force between any two charges /C113and/C1130in a homogeneous isotropic medium is given by Coulomb’s law FˆC/C113/C1130 r2^r; where ris the distance between the charges, and ^ris a unit vector in the direction of the force. The constant /C67determines the system of units, which is not of interest to us; thus we leave /C67as it is. An electric field /C69is said to exist in a region if a stationary charge /C1130in that region experiences a force F: /C69ˆlim /C1130!0…F=/C1130†: The lim /C1130!0guarantees that the test charge /C1130will not alter the charge distribution that existed prior to the introduction of the test charge /C1130. From this definition and Coulomb’s law we find that the electric field at a point rdistant from a point charge is given by /C69ˆC/C113 r2^r: Taking the curl on both sides we get /C114 /C69ˆ0; which shows that the electrostatic field is a conservative field. Hence a potential function /C30exists such that /C69ˆÿ /C114 /C30: Taking the divergence of both sides /C114… /C114 /C30† ˆ ÿ/C114  /C69 390PARTIAL DIFFERENTIAL EQUATIONS or /C1142/C30ˆÿ /C114 /C69: /C114/C69is given by Gauss’ law. To see this, consider a volume /C28containing a total charge /C113. Let dsbe an element of the surface Swhich bounds the volume /C28. Then ZZ S/C69dsˆC/C113ZZ S^rds r2: The quantity ^rdsis the projection of the element area dson a plane perpendi- cular to r. This projected area divided by r2is the solid angle subtended by ds, which is written d/C10. Thus, we have ZZ S/C69dsˆC/C113ZZ S^rds r2ˆC/C113ZZ Sd/C10ˆ4C/C113: If we write /C113as /C113ˆZZZ /C28/C26d/C86; where /C26is the charge density, then ZZ S/C69dsˆ4CZZZ /C28/C26d/C86: But (by the divergence theorem) ZZ S/C69dsˆZZZ /C28/C114/C69d/C86: Substituting this into the previous equation, we obtain ZZZ /C28/C114/C69d/C86ˆ4CZZZ /C28/C26d/C86 or ZZZ /C28/C114/C69ÿ4C/C26 …† d/C86ˆ0: This equation must be valid for all volumes, that is, for any choice of the volume /C28. Thus, we have Gauss’ law in di/C128erential form: /C114/C69ˆ4C/C26: Substituting this into the equation /C1142/C30ˆÿ /C114 /C69, we get /C1142/C30ˆÿ4C/C26; which is Poisson’s equation. In the Gaussian system of units, Cˆ1; in the SI system of units, Cˆ1=4/C340, where the constant /C340is known as the permittivity of free space. If we use SI units, then /C1142/C30ˆÿ/C26=/C34 0: 391LINEAR SECOND-ORDER PDEs In the particular case of zero charge density it reduces to Laplace’s equation, /C1142/C30ˆ0: In the following sections, we shall consider a number of problems to illustrate some useful methods of solving linear partial di/C128erential equations. There are many methods by which homogeneous linear equations with constant coecients can be solved. The following are commonly used in the applications. (1) General solutions: In this method we first find the general solution and then that particular solution which satisfies the boundary conditions. It is always satisfying from the point of view of a mathematician to be able to find general solutions of partial di/C128erential equations; however, general solutions are dicult to find and such solutions are sometimes of little value when given boundary conditions are to be imposed on the solution. To overcome this diculty it is best to find a less general type of solution which is satisfied by the type of boundary conditions to be imposed. This is the method of separation of variables. (2) Separation of variables: The method of separation of variables makes use of the principle of superposition in building up a linear combination of individualsolutions to form a solution satisfying the boundary conditions. The basic approach of this method in attempting to solve a di/C128erential equation (in, say, two dependent variables xandy) is to write the dependent variable u…x;y†as a product of functions of the separate variables u…x;y†ˆX…x†Y…y†. In many cases the partial di/C128erential equation reduces to ordinary di/C128erential equations for X and/C89. (3) Laplace transform method: We first obtain the Laplace transform of the partial di/C128erential equation and the associated boundary conditions with respect to one of the independent variables, and then solve the resulting equation for the Laplace transform of the required solution which can be found by taking the inverse Laplace transform. /C83olutions of Laplace/C39s equation/C58 separation of /C118ariables (1) Laplace’s equation in two dimensions …x;y†: If the potential /C30is a function of only two rectangular coordinates, Laplace’s equation reads /C642/C30 /C64x2‡/C642/C30 /C64y2ˆ0: It is possible to obtain the general solution to this equation by means of a trans- formation to a new set of independent variables: /C24ˆx‡iy;/C17ˆxÿiy; 392PARTIAL DIFFERENTIAL EQUATIONS where Iis the unit imaginary number. In terms of these we have /C64 /C64xˆ/C64 /C64/C24/C64/C24 /C64x‡/C64 /C64/C17/C64/C17 /C64xˆ/C64 /C64/C24‡/C64 /C64/C17; /C642 /C64x2ˆ/C64 /C64x/C64 /C64/C24‡/C64 /C64/C17 ˆ/C64 /C64/C24/C64 /C64/C24‡/C64 /C64/C17/C64/C24 /C64x‡/C64 /C64/C17/C64 /C64/C24‡/C64 /C64/C17/C64/C17 /C64x ˆ/C642 /C64/C242‡2/C64 /C64/C24/C64 /C64/C17‡/C642 /C64/C172: Similarly, we have /C642 /C64y2ˆÿ/C642 /C64/C242‡2/C64 /C64/C24/C64 /C64/C17ÿ/C642 /C64/C172 and Laplace’s equation now reads /C1142/C30ˆ4/C642/C30 /C64/C24/C64/C17ˆ0: Clearly, a very general solution to this equation is /C30ˆf1…/C24†‡f2…/C17†ˆf1…x‡iy†‡f2…xÿiy†; where f1andf2are arbitrary functions which are twice di/C128erentiable. However, it is a somewhat dicult matter to choose the functions f1andf2such that the equation is, for example, satisfied inside a square region defined by the lines xˆ0;xˆa;yˆ0;yˆband such that /C30takes prescribed values on the boundary of this region. For many problems the method of separation of variables is moresatisfactory. Let us apply this method to Laplace’s equation in three dimensions. (2) Laplace’s equation in three dimensions ( x;y;z): Now we have /C642/C30 /C64x2‡/C642/C30 /C64y2‡/C642/C30 /C64z2ˆ0: …10:6† We make the assumption, justifiable by its success, that /C30…x;y;z†may be written as the product /C30…x;y;z†ˆX…x†Y…y†Z…z†: Substitution of this into Eq. (10.6) yields, after division by /C30; 1 Xd2X dx2‡1 Yd2Y dy2ˆÿ1 Zd2Z dz2: …10:7† 393SOLUTIONS OF LAPLACE’S EQUATION The left hand side of Eq. (10.7) is a function of xandy, while the right hand side is a function of zalone. If Eq. (10.7) is to have a solution at all, each side of the equation must be equal to the same constant, say k2 3. Then Eq. (10.7) leads to d2Z dz2‡k23Zˆ0; …10:8† 1 Xd2X dx2ˆÿ1 Yd2Y dy2‡k2 3: …10:9† The left hand side of Eq. (10.9) is a function of xonly, while the right hand side is a function of yonly. Thus, each side of the equation must be equal to a constant, sayk2 1. Therefore d2X dx2‡k21Xˆ0; …10:10† d2Y dy2‡k2 2Yˆ0; …10:11† where k22ˆk21ÿk23: The solution of Eq. (10.10) is of the form X…x†ˆa…k1†ek1x;k16ˆ0;ÿ1 <k1<1 or X…x†ˆa…k1†ek1x‡a0…k1†eÿk1x;k16ˆ0;0<k1<1: …10:12† Similarly, the solutions of Eqs. (10.11) and (10.8) are of the forms Y…y†ˆb…k2†ek2y‡b0…k2†eÿk2y;k26ˆ0;0<k2<1; …10:13† Z…z†ˆc…k3†ek3z‡c0…k3†eÿk3z;k36ˆ0;0<k3<1: …10:14† Hence /C30ˆ‰a…k1†ek1x‡a0…k1†eÿk1xЉb…k2†ek2y‡b0…k2†eÿk2yЉc…k3†ek3z‡c0…k3†eÿk3zŠ; and the general solution of Eq. (10.6) is obtained by integrating the above equa- tion over all the permissible values of the ki…iˆ1;2;3†. In the special case when kiˆ0…iˆ1;2;3†, Eqs. (10.8), (10.10), and (10.11) have solutions of the form Xi…xi†ˆaixi‡bi; where x1ˆx, and X1ˆXetc. 394PARTIAL DIFFERENTIAL EQUATIONS Let us now apply the above result to a simple problem in electrostatics: that of finding the potential /C30at a point Pa distance hfrom a uniformly charged infinite plane in a dielectric of permittivity /C34. Let be the charge per unit area of the plane, and take the origin of the coordinates in the plane and the x-axis perpen- dicular to the plane. It is evident that /C30is a function of xonly. There are two types of solutions, namely: /C30…x†ˆa…k1†ek1x‡a0…k1†eÿk1x; /C30…x†ˆa1x‡b1; the boundary conditions will eliminate the unwanted one. The first boundary condition is that the plane is an equipotential, that is, /C30…0†ˆconstant, and the second condition is that /C69ˆÿ/C64/C30=/C64 xˆ=2/C34. Clearly, only the second type of solution satisfies both the boundary conditions. Hence b1ˆ/C30…0†;a1ˆÿ=2/C34, and the solution is /C30…x†ˆÿ 2/C34x‡/C30…0†: (3) Laplace’s equation in cylindrical coordinates …/C26; ’;z†: The cylindrical co- ordinates are shown in Fig. 10.1, where xˆ/C26cos’ yˆ/C26sin’ zˆz9 >= >;or/C262ˆx2‡y2 ’ˆtanÿ1…y=x† zˆz:8 >< >: Laplace’s equation now reads /C1142/C30…/C26; ’;z†ˆ1 /C26/C64 /C64/C26/C26/C64/C30 /C64/C26 ‡1 /C262/C642/C30 /C64’2‡/C642/C30 /C642z2ˆ0: …10:15† 395SOLUTIONS OF LAPLACE’S EQUATION Figure 10.1. Cylindrical coordinates. We assume that /C30…/C26; ’;z†ˆR…/C26†…’†Z…z†: …10:16† Substitution into Eq. (10.15) yields, after division by /C30, 1 /C26Rd d/C26/C26dR d/C26 ‡1 /C262d2 d’2ˆÿ1 Zd2Z dz2: …10:17† Clearly, both sides of Eq. (10.17) must be equal to a constant, say ÿk2. Then 1 Zd2Z dz2ˆk2ord2Z dz2ÿk2Zˆ0 …10:18† and 1 /C26Rd d/C26/C26dR d/C26 ‡1 /C262d2 d’2ˆÿk2 or /C26 Rd d/C26/C26dR d/C26 ‡k2/C262ˆÿ1 d2 d’2: Both sides of this last equation must be equal to a constant, say 2. Hence d2 d’2‡ 2ˆ0; …10:19† 1 Rd d/C26/C26dR d/C26 ‡k2ÿ 2 /C262/C32! Rˆ0: …10:20† Equation (10.18) has for solutions Z…z†ˆc…k†ekz‡c0…k†eÿkz;k6ˆ0;0<k<1; c1z‡c2; kˆ0;( …10:21† where candc0are arbitrary functions of kandc1andc2are arbitrary constants. Equation (10.19) has solutions of the form …’†ˆa… †ei ’; 6ˆ0;ÿ1 < < 1; b’‡b0; ˆ0:( That the potential must be single-valued requires that …’†ˆ…’‡2n†, where nis an integer. It follows from this that must be an integer or zero and that bˆ0. Then the solution …’†becomes …’†ˆa… †ei ’‡a0… †eÿi ’; 6ˆ0; ˆinteger ; b0; ˆ0:( …10:22† 396PARTIAL DIFFERENTIAL EQUATIONS In the special case kˆ0, Eq. (10.20) has solutions of the form R…/C26†ˆd… †/C26 ‡d0… †/C26ÿ ; 6ˆ0; fln/C26‡/C103; ˆ0:/C26 …10:23† When k6ˆ0, a simple change of variable can put Eq. (10.20) in the form of Bessel’s equation. Let xˆk/C26, then dxˆkd/C26and Eq. (10.20) becomes d2R dx2‡1 xdR dx‡1ÿ 2 x2/C32! Rˆ0; …10:24† the well-known Bessel’s equation (Eq. (7.71)). As shown in Chapter 7, R…x†can be written as R…x†ˆA/C74 …x†‡B/C74ÿ …x†; …10:25† where AandBare constants, and /C74 …x†is the Bessel function of the first kind. When is not an integer, /C74 and/C74ÿ are independent. But when is an integer, /C74ÿ …x†ˆ… ÿ 1†n/C74 …x†, thus /C74 and/C74ÿ are linearly dependent, and Eq. (10.25) cannot be a general solution. In this case the general solution is given by R…x†ˆA1/C74 …x†‡B1Y …x†; …10:26† where A1andB2are constants; Y …x†is the Bessel function of the second kind of order or Neumann’s function of order /C78 …x†. The general solution of Eq. (10.20) when k6ˆ0 is therefore R…/C26†ˆ/C112… †/C74 …k/C26†‡/C113… †Y …k/C26†; …10:27† where pand/C113are arbitrary functions of . Then these functions are also solu- tions: H…1† …k/C26†ˆ/C74 …k/C26†‡iY …k/C26†;H…2† …k/C26†ˆ/C74 …k/C26†ÿiY …k/C26†: These are the Hankel functions of the first and second kinds of order , respec- tively. The functions /C74 ;Y (or/C78 ), and H…1† ,a n d H…2† which satisfy Eq. (10.20) are known as cylindrical functions of integral order and are denoted by Z …k/C26†, which is not the same as Z…z†. The solution of Laplace’s equation (10.15) can now be written /C30…/C26; ’;z†ˆ…c1z‡b†…fln/C26‡/C103†; kˆ0; ˆ0; ˆ…c1z‡b†‰d… †/C26 ‡d0… †/C26ÿ Љa… †ei ’‡a0… †eÿi ’Š; kˆ0; 6ˆ0; ˆ‰c…k†ekz‡c0…k†eÿkzŠZ0…k/C26†; k6ˆ0; ˆ0; ˆ‰c…k†ekz‡c0…k†eÿkzŠZ …k/C26†‰a… †ei ’‡a0… †eÿi ’Š;k6ˆ0; 6ˆ0:8 >>>>>>< >>>>>>: 397SOLUTIONS OF LAPLACE’S EQUATION Let us now apply the solutions of Laplace’s equation in cylindrical coordinates to an infinitely long cylindrical conductor with radius land charge per unit length . We want to find the potential at a point Pa distance /C26/C62/C108from the axis of the cylindrical. Take the origin of the coordinates on the axis of the cylinder that is taken to be the z-axis. The surface of the cylinder is an equipotential: /C30…/C108†ˆconst :forrˆ/C108and all ’andz: The secondary boundary condition is that /C69ˆÿ/C64/C30=/C64/C26 ˆ=2/C108/C34forrˆ/C108and all ’andz: Of the four types of solutions to Laplace’s equation in cylindrical coordinates listed above only the first can satisfy these two boundary conditions. Thus /C30…/C26†ˆb…fln/C26‡/C103†ˆÿ 2/C34ln/C26 /C108‡/C30…a†: (4) Laplace’s equation in spherical coordinates …r; ;’†: The spherical coordinates are shown in Fig. 10.2, where xˆrsincos’; yˆrsinsin’; zˆrcos’: Laplace’s equation now reads /C1142/C30…r; ;’†ˆ1 r/C64 /C64rr2/C64/C30 /C64r ‡1 r2sin/C64 /C64sin/C64/C30 /C64 ‡1 r2sin2/C642/C30 /C64’2ˆ0:…10:28† 398PARTIAL DIFFERENTIAL EQUATIONS Figure 10.2. Spherical coordinates. Again, assume that /C30…r; ;’†ˆR…r†/C2…†…’†: …10:29† Substituting into Eq. (10.28) and dividing by /C30we obtain sin2 Rd drr2dR dr ‡sin /C2d dsind/C2 d ˆÿ1 d2 d’2: For a solution, both sides of this last equation must be equal to a constant, say m2. Then we have two equations d2 d’2‡m2ˆ0; …10:30† sin2 Rd drr2dR dr ‡sin /C2d dsind/C2 d ˆm2; the last equation can be rewritten as 1 /C2sind dsind/C2 d ÿm2 sin2ˆÿ1 Rd drr2dR dr : Again, both sides of the last equation must be equal to a constant, say ÿ/C12. This yields two equations 1 Rd drr2dR dr ˆ/C12; …10:31† 1 /C2sind dsind/C2 d ÿm2 sin2ˆÿ/C12: By a simple substitution: xˆcos, we can put the last equation in a more familiar form: d dx…1ÿx2†dP dx ‡/C12ÿm2 1ÿx2/C32! Pˆ0 …10:32† or …1ÿx2†d2P dx2ÿ2xdP dx‡/C12ÿm2 1ÿx2"# Pˆ0; …10:32a† where we have set P…x†ˆ/C2…†. You may have already noticed that Eq. (10.32) is very similar to Eq. (10.25), the associated Legendre equation. Let us take a close look at this resemblance. In Eq. (10.32), the points xˆ1 are regular singular points of the equation. Let us first study the behavior of the solution near point xˆ1; it is convenient to 399SOLUTIONS OF LAPLACE’S EQUATION bring this regular singular point to the origin, so we make the substitution uˆ1ÿx;U…u†ˆP…x†. Then Eq. (10.32) becomes d duu…2ÿu†dU du ‡/C12ÿm2 u…2ÿu†"# Uˆ0: When we solve this equation by a power series: UˆP1 nˆ0anun‡/C26, we find that the indicial equation leads to the values m=2 for /C26. For the point xˆÿ1, we make the substitution /C118ˆ1‡x, and then solve the resulting di/C128erential equation by the power series method; we find that the indicial equation leads to the same values m=2 for /C26. Let us first consider the value ‡m=2;m0. The above considerations lead us to assume P…x†ˆ… 1ÿx†m=2…1‡x†m=2y…x†ˆ… 1ÿx2†m=2y…x†; m0 as the solution of Eq. (10.32). Substituting this into Eq. (10.32) we find …1ÿx2†d2y dx2ÿ2…m‡1†xdy dx‡/C12ÿm…m‡1† ‰Š yˆ0: Solving this equation by a power series y…x†ˆX1 nˆ0cnxn‡; we find that the indicial equation is …ÿ1†ˆ0. Thus the solution can be written y…x†ˆX nevencnxn‡X noddcnxn: The recursion formula is cn‡2ˆ…n‡m†…n‡m‡1†ÿ/C12 …n‡1†…n‡2†cn: Now consider the convergence of the series. By the ratio test, Rnˆcnxn cnÿ2xnÿ2/C12/C12/C12/C12/C12/C12/C12/C12ˆ …n‡m†…n‡m‡1†ÿ/C12 …n‡1†…n‡2†/C12/C12/C12/C12/C12/C12/C12/C12xjj 2: The series converges for jxj<1, whatever the finite value of /C12may be. For jxjˆ1, the ratio test is inconclusive. However, the integral test yields Z M…t‡m†…t‡m‡1†ÿ/C12 …t‡1†…t‡2†dtˆZ M…t‡m†…t‡m‡1† …t‡1†…t‡2†dtÿZ M/C12 …t‡1†…t‡2†dt and since Z M…t‡m†…t‡m‡1† …t‡1†…t‡2†dt!1 asM!1 ; 400PARTIAL DIFFERENTIAL EQUATIONS the series diverges for jxjˆ1. A solution which converges for all xcan be obtained if either the even or odd series is terminated at the term in xj. This may be done by setting /C12equal to /C12ˆ…j‡m†…j‡m‡1†ˆ/C108…/C108‡1†: On substituting this into Eq. (10.32a), the resulting equation is …1ÿx2†d2P dx2ÿ2xdP dx‡/C108…/C108‡1†ÿm2 1ÿx2"# Pˆ0; which is identical to Eq. (7.25). Special solutions were studied there: they were written in the form Pm /C108…x†and are known as the associated Legendre functions of the first kind of degree land order m, where land m, take on the values /C108ˆ0;1;2;...;andmˆ0;1;2;...;/C108. The general solution of Eq. (10.32) for m0 is therefore P…x†ˆ/C2…†ˆa/C108Pm/C108…x†: …10:33† The second solution of Eq. (10.32) is given by the associated Legendre function of the second kind of degree land order m:/C81m /C108…x†. However, only the associated Legendre function of the first kind remains finite over the range ÿ1x1 (or 02†. Equation (10.31) for R…r†becomes d drr2dR dr ÿ/C108…/C108‡1†Rˆ0: …10:31a† When /C1086ˆ0, its solution is R…r†ˆb…/C108†r/C108‡b0…/C108†rÿ/C108ÿ1; …10:34† and when /C108ˆ0, its solution is R…r†ˆcrÿ1‡d: …10:35† The solution of Eq. (10.30) is ˆf…m†eim’‡f0…/C108†eÿim’;m6ˆ0;positive integer ; /C103; mˆ0:( …10:36† The solution of Laplace’s equation (10.28) is therefore given by /C30…r; ;’†ˆ‰br/C108‡b0rÿ/C108ÿ1ŠPm/C108…cos†‰feim’‡f0eÿim’Š;/C1086ˆ0;m6ˆ0; ‰br/C108‡b0rÿ/C108ÿ1ŠP/C108…cos†; /C1086ˆ0;mˆ0; ‰crÿ1‡dŠP0…cos†; /C108ˆ0;mˆ0;8 >< >:…10:37† where P/C108ˆP0 /C108. 401SOLUTIONS OF LAPLACE’S EQUATION We now illustrate the usefulness of the above result for an electrostatic problem having spherical symmetry. Consider a conducting spherical shell of radius aand charge per unit area. The problem is to find the potential /C30…r; ;’†at a point Pa distance r/C62afrom the center of shell. Take the origin of coordinates to be at the center of the shell. As the surface of the shell is an equipotential, we have the first boundary condition /C30…r†ˆconstant ˆ/C30…a†forrˆaand all and’: …10:38† The second boundary condition is that /C30!0 for r!1 and all and’: …10:39† Of the three types of solutions (10.37) only the last can satisfy the boundaryconditions. Thus /C30…r; ;’†ˆ…cr ÿ1‡d†P0…cos†: …10:40† Now P0…cos†ˆ1, and from Eq. (10.38) we have /C30…a†ˆcaÿ1‡d: But the boundary condition (10.39) requires that dˆ0. Thus /C30…a†ˆcaÿ1,o r cˆa/C30…a†, and Eq. (10.40) reduces to /C30…r†ˆa/C30…a† r: …10:41† Now /C30…a†=aˆ/C69…a†ˆ/C81=4a2/C34; where /C34is the permittivity of the dielectric in which the shell is embedded, /C81ˆ4a2. Thus /C30…a†ˆa=/C34, and Eq. (10.41) becomes /C30…r†ˆa2 /C34r: …10:42† /C83olutions of the /C119a/C118e equation/C58 separation of /C118ariables We now use the method of separation of variables to solve the wave equation /C642u…x;t† /C64x2ˆ/C118ÿ2/C642u…x;t† /C64t2; …10:43† subject to the following boundary conditions: u…0;t†ˆu…/C108;t†ˆ0;t0; …10:44† u…x;0†ˆf…t†;0x/C108; …10:45† 402PARTIAL DIFFERENTIAL EQUATIONS and /C64u…x;t† /C64t/C12/C12/C12/C12 tˆ0ˆ/C103…x†;0x/C108; …10:46† where fandgare given functions. Assuming that the solution of Eq. (10.43) may be written as a product u…x;t†ˆX…x†T…t†; …10:47† then substituting into Eq. (10.43) and dividing by XTwe obtain 1 Xd2X dx2ˆ1 /C1182Td2T dt2: Both sides of this last equation must be equal to a constant, say ÿb2=/C1182. Then we have two equations 1 Xd2X dx2ˆÿb2 /C1182; …10:48† 1 Td2T dt2ˆÿb2: …10:49† The solutions of these equations are periodic, and it is more convenient to write them in terms of trigonometric functions X…x†ˆAsinbx /C118‡Bcosbx /C118; T…t†ˆCsinbt‡Dcosbt; …10:50† where A;B;C, and /C68are arbitrary constants, to be fixed by the boundary condi- tions. Equation (10.47) then becomes u…x;t†ˆ Asinbx /C118‡Bcosbx /C118 …Csinbt‡Dcosbt†: …10:51† The boundary condition u…0;t†ˆ0…t/C620†gives 0ˆB…Csinbt‡Dcosbt† for all t, which implies Bˆ0: …10:52† Next, from the boundary condition u…/C108;t†ˆ0…t/C620†we have 0ˆAsinb/C108 /C118…Csinbt‡Dcosbt†: Note that Bˆ0 would make uˆ0. However, the last equation can be satisfied for all twhen sinb/C108 /C118ˆ0; 403SOLUTIONS OF LAPLACE’S EQUATION which implies bˆn/C118 /C108;nˆ1;2;3;...: …10:53† Note that ncannot be equal to zero, because it would make bˆ0, which in turn would make uˆ0. Substituting Eq. (10.53) into Eq. (10.51) we have un…x;t†ˆsinnx /C108Cnsinn/C118t /C108‡Dncosn/C118t /C108 ; nˆ1;2;3;...:…10:54† We see that there is an infinite set of discrete values of band that to each value of bthere corresponds a particular solution. Any linear combination of these parti- cular solutions is also a solution: un…x;t†ˆX1 nˆ1sinnx /C108Cnsinn/C118t /C108‡Dncosn/C118t /C108 : …10:55† The constants CnandDnare fixed by the boundary conditions (10.45) and (10.46). Application of boundary condition (10.45) yields f…x†ˆX1 nˆ1Dnsinnx /C108: …10:56† Similarly, application of boundary condition (10.46) gives /C103…x†ˆ/C118 /C108X1 nˆ1nCnsinnx /C108: …10:57† The coecients CnandDnmay then be determined by the Fourier series method: Dnˆ2 /C108Z/C108 0f…x†sinnx /C108dx; Cnˆ2 n/C118Z/C108 0/C103…x†sinnx /C108dx: …10:58† We can use the method of separation of variable to solve the heat conduction equation. We shall leave this as a home work problem. In the following sections, we shall consider two more methods for the solution of linear partial di/C128erential equations: the method of Green’s functions, and the method of the Laplace transformation which was used in Chapter 9 for the solution of ordinary linear di/C128erential equations with constant coecients. /C83olution of Poisson/C39s equation/C46 /C71reen/C39s functions The Green’s function approach to boundary-value problems is a very powerfultechnique. The field at a point caused by a source can be considered to be the total e/C128ect due to each ‘‘unit’’ (or elementary portion) of the source. If /C71…x;x 0†is the 404PARTIAL DIFFERENTIAL EQUATIONS field at a point xdue to a unit point source at x0, then the total field at xdue to a distributed source /C26…x0†is the integral of /C71/C26over the range of x0occupied by the source. The function /C71…x;x0†is the well-known Green’s function. We now apply this technique to solve Poisson’s equation for electric potential /C30(Example 10.2) /C1142/C30…r†ˆÿ1 /C34/C26…r†; …10:59† where /C26is the charge density and /C34the permittivity of the medium, both are given. By definition, Green’s function /C71…r;r0†is the solution of /C1142/C71…r;r0†ˆ…rÿr0†; …10:60† where …rÿr0†is the Dirac delta function. Now, multiplying Eq. (10.60) by /C30and Eq. (10.59) by /C71, and then subtracting, we find /C30…r†/C1142/C71…r;r0†ÿ/C71…r;r0†/C1142/C30…r†ˆ/C30…r†…rÿr0†‡1 /C34/C71…r;r0†/C26…r†; and on interchanging rand r0, /C30…r0†/C11402/C71…r0;r†ÿ/C71…r0;r†/C11402/C30…r0†ˆ/C30…r0†…r0ÿr†‡1 /C34/C71…r0;r†/C26…r0† or /C30…r0†…r0ÿr†ˆ/C30…r0†/C11402/C71…r0;r†ÿ/C71…r0;r†/C11402/C30…r0†ÿ1 /C34/C71…r0;r†/C26…r0†;…10:61† the prime on /C114indicates that di/C128erentiation is with respect to the primed co- ordinates. Integrating this last equation over all r0within and on the surface S0 which encloses all sources (charges) yields /C30…r†ˆÿ1 /C34Z /C71…r;r0†/C26…r0†dr0 ‡Z ‰/C30…r0†/C11402/C71…r;r0†ÿ/C71…r;r0†/C11402/C30…r0†Šdr0; …10:62† where we have used the property of the delta function Z‡1 ÿ1f…r0†…rÿr0†dr0ˆf…r†: We now use Green’s theorem ZZZ f/C11402/C32ÿ/C32/C11402f†d/C280ˆZZ …f/C1140/C32ÿ/C32/C1140f†d/C83 405SOLUTIONS OF POISSON’S EQUATION to transform the second term on the right hand side of Eq. (10.62) and obtain /C30…r†ˆÿ1 /C34Z /C71…r;r0†/C26…r0†dr0 ‡Z ‰/C30…r0†/C1140/C71…r;r0†ÿ/C71…r;r0†/C1140/C30…r0†Š d/C830…10:63† or /C30…r†ˆÿ1 /C34Z /C71…r;r0†/C26…r0†dr0 ‡Z /C30…r0†/C64 /C64n0/C71…r;r0†ÿ/C71…r;r0†/C64 /C64n0/C30…r0† d/C830; …10:64† where n0is the outward normal to dS0. The Green’s function /C71…r;r0†can be found from Eq. (10.60) subject to the appropriate boundary conditions. If the potential /C30vanishes on the surface S0or/C64/C30=/C64n0vanishes, Eq. (10.64) reduces to /C30…r†ˆÿ1 /C34Z /C71…r;r0†/C26…r0†dr0: …10:65† On the other hand, if the surface S0encloses no charge, then Poisson’s equation reduces to Laplace’s equation and Eq. (10.64) reduces to /C30…r†ˆZ /C30…r0†/C64 /C64n0/C71…r;r0†ÿ/C71…r;r0†/C64 /C64n0/C30…r0† d/C830: …10:66† The potential at a field point rdue to a point charge /C113located at the point r0is /C30…r†ˆ1 4/C34/C113 rÿr0jj: Now /C1142 1 rÿr0jj ˆÿ4…rÿr0† (the proof is left as an exercise for the reader) and it follows that the Green’s function /C71…r;r0†in this case is equal /C71…r;r0†ˆ1 4/C341 rÿr0jj: If the medium is bounded, the Green’s function can be obtained by direct solutionof Eq. (10.60) subject to the appropriate boundary conditions. To illustrate the procedure of the Green’s function technique, let us consider a simple example that can easily be solved by other methods. Consider two grounded parallel conducting plates of infinite extent: if the electric charge density /C26between the two plates is given, find the electric potential distribution /C30between 406PARTIAL DIFFERENTIAL EQUATIONS the plates. The electric potential distribution /C30is described by solving Poisson’s equation /C1142/C30ˆÿ/C26=/C34 subject to the boundary conditions (1)/C30…0†ˆ0; (2)/C30…1†ˆ0: We take the coordinates shown in Fig. 10.3. Poisson’s equation reduces to the simple form d2/C30 dx2ˆÿ/C26 /C34: …10:67† Instead of using the general result (10.64), it is more convenient to proceed directly. Multiplying Eq. (10.67) by /C71…x;x0†and integrating, we obtain Z1 0/C71d2/C30 dx2dxˆÿZ1 0/C26…x†/C71 /C34dx: …10:68† Then using integration by parts gives Z1 0/C71d2/C30 dx2dxˆ/C71…x;x0†d/C30…x† dx/C12/C12/C12/C121 0ÿZ1 0d/C71 dxd/C30 dxdx and using integration by parts again on the right hand side, we obtain ÿZ1 0/C71d2/C30 dx2dxˆÿ/C71…x;x0†d/C30…x† dx1 0‡d/C71 dx/C3010 ÿZ1 0/C30d2/C71 dx2dx/C12/C12/C12/C12/C12"#/C12/C12/C12/C12/C12 ˆ/C71…0;x 0†d/C30…0† dxÿ/C71…1;x0†d/C30…1† dxÿZ1 0/C30d2/C71 dx2dx: 407SOLUTIONS OF POISSON’S EQUATION Figure 10.3. Substituting this into Eq. (10.68) we obtain /C71…0;x0†d/C30…0† dxÿ/C71…1;x0†d/C30…1† dxÿZ1 0/C30d2/C71 dx2dxˆZ1 0/C71…x;x0†/C26…x† /C34dx or Z1 0/C30d2/C71 dx2dxˆ/C71…1;x0†d/C30…1† dxÿ/C71…0;x0†d/C30…0† dxÿZ1 0/C71…x;x0†/C26…x† /C34dx:…10:69† We must now choose a Green’s function which satisfies the following equation and the boundary conditions: d2/C71 dx2ˆÿ…xÿx0†;/C71…0;x0†ˆ/C71…1;x0†ˆ0: …10:70† Combining these with Eq. (10.69) we find the solution to be /C30…x0†ˆZ1 01 /C34/C26…x†/C71…x;x0†dx: …10:71† It remains to find /C71…x;x0†. By integration, we obtain from Eq. (10.70) d/C71 dxˆÿZ …xÿx0†dx‡aˆÿU…xÿx0†‡a; where Uis the unit step function and ais an integration constant to be determined later. Integrating once we get /C71…x;x0†ˆÿZ U…xÿx0†dx‡ax‡bˆÿ … xÿx0†U…xÿx0†‡ax‡b: Imposing the boundary conditions on this general solution yields two equations: /C71…0;x0†ˆx0U…ÿx0†‡a0‡bˆ0‡0‡bˆ0; /C71…1;x0†ˆÿ … 1ÿx0†U…1ÿx0†‡a‡bˆ0: From these we find aˆ…1ÿx0†U…1ÿx0†;bˆ0 and the Green’s function is /C71…x;x0†ˆÿ … xÿx0†U…xÿx0†‡…1ÿx0†x: …10:72† This gives the response at x0due to a unit source at x. Interchanging xandx0in Eqs. (10.70) and (10.71) we find the solution of Eq. (10.67) to be /C30…x†ˆZ1 01 /C34/C26…x0†/C71…x0;x†dx0ˆZ1 01 /C34/C26…x0†‰ÿ…x0ÿx†U…x0ÿx†‡…1ÿx†x0Šdx0: …10:73† 408PARTIAL DIFFERENTIAL EQUATIONS Note that the Green’s function in the last equation can be written in the form /C71…x;x0†ˆ…1ÿx†xx <x0 …1ÿx†x0x/C62x0( : Laplace transform solutions of boundar/C121-/C118alue problems Laplace and Fourier transforms are useful in solving a variety of partial di/C128er- ential equations, the choice of the appropriate transforms depends on the type of boundary conditions imposed on the problem. To illustrate the use of the Lapace transforms in solving boundary-value problems, we solve the following equation: /C64u /C64tˆ2/C642u /C64x2; …10:74† u…0;t†ˆu…3;t†ˆ0; u…x;0†ˆ10 sin 2 xÿ6 sin 4 x: …10:75† Taking the Laplace transform of Eq. (10.74) with respect to tgives L/C64u /C64t ˆ2L/C642u /C64x2"# : Now L/C64u /C64t ˆ/C112L u…† ÿ u…x;0† and L/C642u /C64x2"# ˆZ1 0eÿ/C112t/C642u /C64x2dtˆ/C642 /C64x2Z1 0eÿ/C112tu…x;t†dtˆ/C642 /C64x2L‰uŠ: Here /C642=/C64x2andR1 0dtare interchangeable because xandtare independent. For convenience, let UˆU…x;/C112†ˆL‰u…x;t†Š ˆZ1 0eÿ/C112tu…x;t†dt: We then have /C112Uÿu…x;0†ˆ2d2U dx2; from which we obtain, on using the given condition (10.75), d2U dx2ÿ1 2/C112Uˆ3 sin 4 xÿ5 sin 2 x: …10:76† 409BOUNDARY-VALUE PROBLEMS Now think of this as a di/C128erential equation in terms of x,w i t h pas a parameter. Then taking the Laplace transform of the given conditions u…0;t†ˆu…3;t†ˆ0, we have L‰u…0;t†Š ˆ0;L‰u…3;t†Š ˆ0 or U…0;/C112†ˆ0;U…3;/C112†ˆ0: These are the boundary conditions on U…x;/C112†. Solving Eq. (10.76) subject to these conditions we find U…x;/C112†ˆ5 sin 2 x /C112‡162ÿ3 sin 4 x /C112‡642: The solution to Eq. (10.74) can now be obtained by taking the inverse Laplace transform u…x;t†ˆLÿ1‰U…x;/C112†Š ˆ5eÿ162tsin 2xÿ3eÿ642sin 4x: The Fourier transform method was used in Chapter 4 for the solution of ordinary linear ordinary di/C128erential equations with constant coecients. It can be extended to solve a variety of partial di/C128erential equations. However, we shall not discuss this here. Also, there are other methods for the solution of linear partial di/C128erential equations. In general, it is a dicult task to solve partial di/C128erential equations analytically, and very often a numerical method is the best way of obtaining a solution that satisfies given boundary conditions. Problems 10.1 ( a) Show that y…x;t†ˆ/C70…2x‡5t†‡/C71…2xÿ5t†is a general solution of 4/C642y /C64t2ˆ25/C642y /C64x2: (b) Find a particular solution satisfying the conditions y…0;t†ˆy…;t†ˆ0;y…x;0†ˆsin 2x;y0…x;0†ˆ0: 10.2. State the nature of each of the following equations (that is, whether elliptic, parabolic, or hyperbolic) …a†/C642y /C64t2‡ /C642y /C64x2ˆ0; …b†x/C642u /C64x2‡y/C642u /C64y2‡3y2/C64u /C64x: 10.3 The electromagnetic wave equation: Classical electromagnetic theory was worked out experimentally in bits and pieces by Coulomb, Oersted, Ampere, Faraday and many others, but the man who put it all together and built it into the compact and consistent theory it is today was James Clerk Maxwell. 410PARTIAL DIFFERENTIAL EQUATIONS His work led to the understanding of electromagnetic radiation, of which light is a special case. Given the four Maxwell equations /C114/C69ˆ/C26=/C34 0; …Gauss’ law †; /C114 /C66ˆ/C220/C106‡/C340/C64/C69=/C64t …† … Ampere’s law †; /C114/C66ˆ0 …Gauss’ law †; /C114 /C69ˆÿ/C64/C66=/C64t …Faraday’s law †; where /C66is the magnetic induction, /C106ˆ/C26/C118is the current density, and /C220is the permeability of the medium, show that: (a) the electric field and the magnetic induction can be expressed as /C69ˆÿ /C114 /C30ÿ/C64/C65=/C64t; /C66ˆ/C114 /C65; where /C65is called the vector potential, and /C30the scalar potential. It should be noted that /C69and /C66are invariant under the following trans- formations: /C650ˆ/C65‡/C114/C31; /C300ˆ/C30ÿ/C64/C30=/C64 t in which /C31is an arbitrary real function. That is, both ( /C650;/C30†, and (/C650;/C300) yield the same /C69and/C66. Any condition which, for computational convenience, restricts the form of /C65and/C30is said to define a gauge. Thus the above transformation is called a gauge transformation and /C31is called a gauge parameter. (b) If we impose the so-called Lorentz gauge condition on /C65and/C30: /C114/C65‡/C220/C340…/C64/C30=/C64 t†ˆ0; then both /C65and/C30satisfy the following wave equations: /C1142/C65ÿ/C220/C340/C642/C65 /C64t2ˆÿ/C220/C106; /C1142/C30ÿ/C220/C340/C642/C30 /C64t2ˆÿ/C26=/C34 0: 10.4 Given Gauss’ lawRR S/C69dsˆ/C113=/C34, find the electric field produced by a charged plane of infinite extension is given by /C69ˆ=/C34, where is the charge per unit area of the plane. 10.5 Consider an infinitely long uncharged conducting cylinder of radius lplaced in an originally uniform electric field /C690directed at right angles to the axis of the cylinder. Find the potential at a point /C26…/C62/C108†from the axis of the cylin- der. The boundary conditions are: 411PROBLEMS /C30…/C26; ’†ˆÿ/C690/C26cos’ˆÿ/C690xfor /C26!1 ; 0f o r /C26ˆ/C108;/C26 where the x-axis has been taken in the direction of the uniform field /C690. 10.6 Obtain the solution of the heat conduction equation /C642u…x;t† /C64x2ˆ1 /C64u…x;t† /C64t which satisfies the boundary conditions (1)u…0;t†ˆu…/C108;t†ˆ0;t0;(2)u…x;0†ˆf…x†;0x, where f…x†is a given function and lis a constant. 10.7 If a battery is connected to the plates as shown in Fig. 10.4, and if the charge density distribution between the two plates is still given by /C26…x†, find the potential distribution between the plates. 10.8 Find the Green’s function that satisfies the equation d2/C71 dx2ˆ…xÿx0† and the boundary conditions /C71ˆ0 when xˆ0 and /C71remains bounded when xapproaches infinity. (This Green’s function is the potential due to a surface charge ÿ/C34per unit area on a plane of infinite extent located at xˆx0 in a dielectric medium of permittivity /C34when a grounded conducting plane of infinite extent is located at xˆ0.) 10.9 Solve by Laplace transforms the boundary-value problem /C642u /C64x2ˆ1 /C75/C64u /C64tfor x/C620;t/C620; given that uˆu0(a constant) on xˆ0 for t/C620, and uˆ0 for x/C620;tˆ0. 412PARTIAL DIFFERENTIAL EQUATIONS Figure 10.4. 11 Simple linear integral e/C113uations In previous chapters we have met equations in which the unknown functions appear under an integral sign. Such equations are called integral equations. Fourier and Laplace transforms are important integral equations, In Chapter 4, by introducing the method of Green’s function we were led in a natural way to reformulate the problem in terms of integral equations. Integral equations have become one of the very useful and sometimes indispensable mathematical tools of theoretical physics and engineering. /C67lassification of linear integral equations In this chapter we shall confine our attention to linear integral equations. Linearintegral equations can be divided into two major groups: (1) If the unknown function occurs only under the integral sign, the integral equation is said to be of the first kind. Integral equations having the unknown function both inside and outside the integral sign are of the second kind. (2) If the limits of integration are constants, the equation is called a Fredholm integral equation. If one limit is variable, it is a Volterra equation. These four kinds of linear integral equations can be written as follows: f…x†ˆZ b a/C75…x;t†u…t†dt Fredholm equation of the first kind; …11:1† u…x†ˆf…x†‡Zb a/C75…x;t†u…t†dtFredholm equation of the second kind; …11:2† 413 f…x†ˆZx a/C75…x;t†u…t†dt Volterra equation of the first kind; …11:3† u…x†ˆf…x†‡Zx a/C75…x;t†u…t†dtVolterra equation of the second kind : …11:4† In each case u…t†is the unknown function, /C75…x;t†andf…x†are assumed to be known. /C75…x;t†is called the kernel or nucleus of the integral equation. is a parameter, which often plays the role of an eigenvalue. The equation is said to be homogeneous if f…x†ˆ0. If one or both of the limits of integration are infinite, or the kernel /C75…x;t† becomes infinite in the range of integration, the equation is said to be singular; special techniques are required for its solution. The general linear integral equation may be written as /C104…x†u…x†ˆf…x†‡Zb a/C75…x;t†u…t†dt: …11:5† If/C104…x†ˆ0, we have a Fredholm equation of the first kind; if /C104…x†ˆ1, we have a Fredholm equation of the second kind. We have a Volterra equation when theupper limit is x. It is beyond the scope of this book to present the purely mathematical general theory of these various types of equations. After a general discussion of a fewmethods of solution, we will illustrate them with some simple examples. We will then show with a few examples from physical problems how to convert di/C128erential equations into integral equations. /C83ome methods of solution Separable /C107ernel When the two variables xand twhich appear in the kernel /C75…x;t†are separable, the problem of solving a Fredholm equation can be reduced to that of solving a system of algebraic equations, a much easier task. When the kernel /C75…x;t†can be written as /C75…x;t†ˆX n iˆ1/C103i…x†/C104i…t†; …11:6† where /C103…x†is a function of xonly and /C104…t†a function of tonly, it is said to be degenerate. Putting Eq. (11.6) into Eq. (11.2), we obtain u…x†ˆf…x†‡Xn iˆ1Zb a/C103i…x†/C104i…t†u…t†dt: 414SIMPLE LINEAR INTEGRAL EQUATIONS Note that /C103…x†is a constant as far as the tintegration is concerned, hence it may be taken outside the integral sign and we have u…x†ˆf…x†‡Xn iˆ1/C103i…x†Zb a/C104i…t†u…t†dt: …11:7† Now Zb a/C104i…t†u…t†dtˆCi…ˆconst :†: …11:8† Substituting this into Eq. (11.7) and solving for u…t†, we obtain u…t†ˆf…x†‡CXn iˆ1/C103i…x†: …11:9† The value of Cimay now be obtained by substituting Eq. (11.9) into Eq. (11.8). The solution is only valid for certain values of , and we call these the eigenvalues of the integral equation. The homogeneous equation has non-trivial solutions only if is one of these eigenvalues; these solutions are called eigenfunctions of the kernel (operator) /C75. Example 11.1As an example of this method, we consider the following equation: u…x†ˆx‡Z 1 0…xt2‡x2t†u…t†dt: …11:10† This is a Fredholm equation of the second kind, with f…x†ˆxand /C75…x;t†ˆxt2‡x2t. If we define ˆZ1 0t2u…t†dt;/C12 ˆZ1 0tu…t†dt; …11:11† then Eq. (11.10) becomes u…x†ˆx‡… x‡/C12x2†: …11:12† To determine Aand B, we put Eq. (11.12) back into Eq. (11.11) and obtain ˆ1 4‡14 ‡15/C12; /C12 ˆ13‡13 ‡14/C12: …11:13† Solving this for and/C12we find ˆ60‡ 240ÿ120ÿ2;/C12 ˆ80 240ÿ120ÿ2; and the final solution is u…t†ˆ…240ÿ60†x‡80x2 240ÿ120ÿ2: 415SOME METHODS OF SOLUTION The solution blows up when ˆ117:96 or ˆ2:04. These are the eigenvalues of the integral equation. Fredholm found that if: (1) f…x†is continuous, (2) /C75…x;t†is piecewise contin- uous, (3) the integralsRR /C752…x;t†dxdt;R f2…t†dtexist, and (4) the integralsRR /C752…x;t†dtandR /C752…t;x†dtare bounded, then the following theorems apply: (a) Either the inhomogeneous equation u…x†ˆf…x†‡Zb a/C75…x;t†u…t†dt has a unique solution for any function f…x†…is not an eigenvalue), or the homogeneous equation u…x†ˆZb a/C75…x;t†u…t†dt has at least one non-trivial solution corresponding to a particular value of . In this case, is an eigenvalue and the solution is an eigenfunction. (b)I fis an eigenvalue, then is also an eigenvalue of the transposed equation u…x†ˆZb a/C75…t;x†u…t†dt; and, if is not an eigenvalue, then is also not an eigenvalue of the transposed equation u…x†ˆf…x†‡Zb a/C75…t;x†u…t†dt: (c)I fis an eigenvalue, the inhomogeneous equation has a solution if, and only if, Zb au…x†f…x†dxˆ0 for every function f…x†. We refer the readers who are interested in the proof of these theorems to the book by R. Courant and D. Hilbert ( Methods of Mathematical Physics , Vol. 1, Wiley, 1961). Neumann series solutions This method is due largely to Neumann, Liouville, and Volterra. In this method we solve the Fredholm equation (11.2) u…x†ˆf…x†‡Zb a/C75…x;t†u…t†dt 416SIMPLE LINEAR INTEGRAL EQUATIONS by iteration or successive approximations, and begin with the approximation u…x†/C25u0…x†/C25f…x†: This approximation is equivalent to saying that the constant or the integral is small. We then put this crude choice into the integral equation (11.2) under the integral sign to obtain a second approximation: u1…x†ˆf…x†‡Zb a/C75…x;t†f…t†dt and the process is then repeated and we obtain u2…x†ˆf…x†‡Zb a/C75…x;t†f…t†dt‡2Zb aZb a/C75…x;t†/C75…t;t0†f…t0†dt0dt: We can continue iterating this process, and the resulting series is known as theNeumann series, or Neumann solution: u…x†ˆf…x†‡Z b a/C75…x;t†f…t†dt‡2Zb aZb a/C75…x;t†/C75…t;t0†f…t0†dt0dt‡ : This series can be written formally as un…x†ˆXn iˆ1i’i…x†; …11:14† where ’0…x†ˆu0…x†ˆf…x†; ’1…x†ˆZb a/C75…x;t1†f…t1†dt1; ’2…x†ˆZb aZb a/C75…x;t1†/C75…t1;t2†f…t2†dt1dt2; ... ’n…x†ˆZb aZb aZb a/C75…x;t1†/C75…t1;t2†/C75…tnÿ1;tn†f…tn†dt1dt2dtn:9 >>>>>>>>>>>>>>= >>>>>>>>>>>>>>; …11:15† The series (11.14) will converge for suciently small , when the kernel /C75…x;t†is bounded. This can be checked with the Cauchy ratio test (Problem 11.4). Example 11.2 Use the Neumann method to solve the integral equation u…x†ˆf…x†‡ 1 2Z1 ÿ1/C75…x;t†u…t†dt; …11:16† 417SOME METHODS OF SOLUTION where f…x†ˆx; /C75…x;t†ˆtÿx: Solution: We begin with u0…x†ˆf…x†ˆx: Then u1…x†ˆx‡1 2Z1 ÿ1…tÿx†tdtˆx‡13: Putting u 1…x†into Eq. (11.16) under the integral sign, we obtain u2…x†ˆx‡12Z 1 ÿ1…tÿx†t‡13 dtˆx‡13ÿx 3: Repeating this process of substituting back into Eq. (11.16) once more, we obtain u3…x†ˆx‡13ÿx 3ÿ1 32: We can improve the approximation by iterating the process, and the convergence of the resulting series (solution) can be checked out with the ratio test. The Neumann method is also applicable to the Volterra equation, as shown by the following example. Example 11.3 Use the Neumann method to solve the Volterra equation u…x†ˆ1‡Zx 0u…t†dt: Solution: We begin with the zeroth approximation u0…x†ˆ1. Then u1…x†ˆ1‡Zx 0u0…t†dtˆ1‡Zx 0dtˆ1‡x: This gives u2…x†ˆ1‡Zx 0u1…t†dtˆ1‡Zx 0…1‡t†dtˆ1‡x‡1 22x2; similarly, u3…x†ˆ1‡Zx 01‡t‡12 2t2 dtˆ1‡t‡12 2t2‡1 3/C333x3: 418SIMPLE LINEAR INTEGRAL EQUATIONS By induction un…x†ˆXn kˆ11 k/C33kxk: When n!1 ,un…x†approaches u…x†ˆex: /C84ransformation of an integral equation into a di/C128erential equation Sometimes the Volterra integral equation can be transformed into an ordinary di/C128erential equation which may be easier to solve than the original integral equa- tion, as shown by the following example. Example 11.4 Consider the Volterra integral equation u…x†ˆ2x‡4Rx 0…tÿx†u…t†dt. Before we transform it into a di/C128erential equation, let us recall the following very usefulformula: if I… †ˆZ b… † a… †f…x; †dx; where aandbare continuous and at least once di/C128erentiable functions of , then dI… † d ˆf…b; †db d ÿf…a; †da d ‡Zb a/C64f…x; † /C64 dx: With the help of this formula, we obtain d dxu…x†ˆ2‡4…tÿx†u…t† fgtˆxÿZx 0u…t†dt ˆ2ÿ4Zx 0u…t†dt: Di/C128erentiating again we obtain d2u…x† dx2ˆÿ4u…x†: This is a di/C128erentiation equation equivalent to the original integral equation, butits solution is much easier to find: u…x†ˆAcos 2 x‡Bsin 2x; where Aand Bare integration constants. To determine their values, we put the solution back into the original integral equation under the integral sign, and then 419SOME METHODS OF SOLUTION integration gives Aˆ0 and Bˆ1. Thus the solution of the original integral equation is u…x†ˆsin 2x: /C76aplace transform solution The Volterra integral equation can sometime be solved with the help of the Laplace transformation and the convolution theorem. Before we consider the Laplace transform solution, let us review the convolution theorem. If f1…x†and f2…x†are two arbitrary functions, we define their convolution ( faltung in German) to be /C103…x†ˆZ1 ÿ1f1…y†f2…xÿy†dy: Its Laplace transform is L‰/C103…x†Š ˆL‰f1…x†ŠL‰f2…x†Š: We now consider the Volterra equation u…x†ˆf…x†‡Zx 0/C75…x;t†u…t†dt ˆf…x†‡Zx 0/C103…xÿt†u…t†dt; …11:17† where /C75…xÿt†ˆ/C103…xÿt†, a so-called displacement kernel. Taking the Laplace transformation and using the convolution theorem, we obtain LZx 0/C103…xÿt†u…t†dt ˆL/C103…xÿt† ‰Š Lu…t†‰Š ˆ /C71…/C112†U…/C112†; where U…/C112†ˆL‰u…t†Š ˆR1 0eÿ/C112tu…t†dt, and similarly for /C71…/C112†. Thus, taking the Laplace transformation of Eq. (11.17), we obtain U…/C112†ˆ/C70…/C112†‡/C71…/C112†U…/C112† or U…/C112†ˆ/C70…/C112† 1ÿ/C71…/C112†: Inverting this we obtain u…t†: u…t†ˆLÿ1 /C70…/C112† 1ÿ/C71…/C112† : 420SIMPLE LINEAR INTEGRAL EQUATIONS Fourier transform solution If the kernel is a displacement kernel and if the limits are ÿ1and‡1, we can use Fourier transforms. Consider a Fredholm equation of the second kind u…x†ˆf…x†‡Z1 ÿ1/C75…xÿt†u…t†dt: …11:18† Taking Fourier transforms (indicated by overbars) 1  2pZ1 ÿ1dxf…x†eÿi/C112xˆf…/C112†;etc:; and using the convolution theorem Z1 ÿ1f…t†/C103…xÿt†dtˆZ1 ÿ1f…y†/C103…y†eÿiyxdy; we obtain the transform of our integral equation (11.18): u…/C112†ˆ f…/C112†‡/C75…/C112†u…/C112†: Solving for u…/C112†we obtain u…/C112†ˆf…/C112† 1ÿ/C75…/C112†: If we can invert this equation, we can solve the original integral equation: u…x†ˆ1 2pZ 1 ÿ1f…t†eÿixt 1ÿ 2p /C75…t†: …11:19† /C84he /C83chmidt/C177/C72ilbert method of solution In many physical problems, the kernel may be symmetric. In such cases, the integral equation may be solved by a method quite di/C128erent from any of thosein the preceding section. This method, devised by Schmidt and Hilbert, is based on considering the eigenfunctions and eigenvalues of the homogeneous integral equation. A kernel /C75…x;t†is said to be symmetric if /C75…x;t†ˆ/C75…t;x†and Hermitian if /C75…x;y†ˆ/C75/C42…t;x†. We shall limit our discussion to such kernels. (a) The homogeneous Fredholm equation u…x†ˆZ b a/C75…x;t†u…t†dt: 421THE SCHMIDT–HILBERT METHOD OF SOLUTION A Hermitian kernel has at least one eigenvalue and it may have an infinite num- ber. The proof will be omitted and we refer interested readers to the book byCourant and Hibert mentioned earlier (Chapter 3). The eigenvalues of a Hermitian kernel are real, and eigenfunctions belonging to di/C128erent eigenvalues are orthogonal; two functions f…x†and/C103…x†are said to be orthogonal if Z f/C42…x†/C103…x†dxˆ0: To prove the reality of the eigenvalue, we multiply the homogeneous Fredholm equation by u/C42…x†, then integrating with respect to x, we obtain Z b au/C42…x†u…x†dxˆZb aZb a/C75…x;t†u/C42…x†u…t†dtdx: …11:20† Now, multiplying the complex conjugate of the Fredholm equation by u…x†and then integrating with respect to x,w eg e t Zb au/C42…x†u…x†dxˆ/C42Zb aZb a/C75/C42…x;t†u/C42…t†u…x†dtdx: Interchanging xandton the right hand side of the last equation and remembering that the kernel is Hermitian /C75/C42…t;x†ˆ/C75…x;t†, we obtain Zb au/C42…x†u…x†dxˆ/C42Zb aZb a/C75…x;t†u…t†u/C42…x†dtdx: Comparing this equation with Eq. (11.2), we see that ˆ/C42, that is, is real. We now prove the orthogonality. Let i,jbe two di/C128erent eigenvalues and ui…x†;uj…x†, the corresponding eigenfunctions. Then we have ui…x†ˆiZb a/C75…x;t†ui…t†dt; uj…x†ˆjZb a/C75…x;t†uj…t†dt: Now multiplying the first equation by uj…x†, the second by iui…x†, and then integrating with respect to x, we obtain jZb aui…x†uj…x†dxˆijZb aZb a/C75…x;t†ui…t†uj…x†dtdx; iZb aui…x†uj…x†dxˆijZb aZb a/C75…x;t†uj…t†ui…x†dtdx:…11:21† Now we interchange xandton the right hand side of the last integral and because of the symmetry of the kernel, we have iZb aui…x†uj…x†dxˆijZb aZb a/C75…x;t†ui…t†uj…x†dtdx: …11:22† 422SIMPLE LINEAR INTEGRAL EQUATIONS Subtracting Eq. (11.21) from Eq. (11.22), we obtain …iÿj†Zb aui…x†uj…x†dxˆ0: …11:23† Since i6ˆj, it follows that Zb aui…x†uj…x†dxˆ0: …11:24† Such functions may always be nomalized. We will assume that this has been done and so the solutions of the homogeneous Fredholm equation form a complete orthonomal set: Zb aui…x†uj…x†dxˆij: …11:25† Arbitrary functions of x, including the kernel for fixed t, may be expanded in terms of the eigenfunctions /C75…x;t†ˆX Ciui…x†: …11:26† Now substituting Eq. (11.26) into the original Fredholm equation, we have uj…t†ˆjZb a/C75…t;x†uj…x†dxˆjZb a/C75…x;t†uj…x†dx ˆjX iZb aCiui…x†uj…x†dxˆjX iCiijˆjCj or Ciˆui…t†=i and for our homogeneous Fredholm equation of the second kind the kernel maybe expressed in terms of the eigenfunctions and eigenvalues as /C75…x;t†ˆX 1 nˆ1un…x†un…t† n: …11:27† The Schmidt–Hilbert theory does not solve the homogeneous integral equation; its main function is to establish the properties of the eigenvalues (reality) and eigenfunctions (orthogonality and completeness). The solutions of the homoge- neous integral equation come from the preceding section on methods of solution. (b) Solution of the inhomogeneous equation u…x†ˆf…x†‡Zb a/C75…x;t†u…t†dt: …11:28† We assume that we have found the eigenfunctions of the homogeneous equation by the methods of the preceding section, and we denote them by ui…x†. We may 423THE SCHMIDT–HILBERT METHOD OF SOLUTION now expand both u…x†andf…x†in terms of ui…x†, which forms an orthonormal complete set. u…x†ˆX1 nˆ1 nun…x†;f…x†ˆX1 nˆ1/C12nun…x†: …11:29† Substituting Eq. (11.29) into Eq. (11.28), we obtain Xn nˆ1 nun…x†ˆXn nˆ1/C12nun…x†‡Zb a/C75…x;t†Xn nˆ1 nun…t†dt ˆXn nˆ1/C12nun…x†‡X1 nˆ1 num…x† mZb aum…t†un…t†dt; ˆXn nˆ1/C12nun…x†‡X1 nˆ1 num…x† mnm; from which it follows that Xn nˆ1 nun…x†ˆXn nˆ1/C12nun…x†‡X1 nˆ1 nun…x† n: …11:30† Multiplying by ui…x†and then integrating with respect to xfrom atob, we obtain nˆ/C12n‡ n=n; …11:31† which can be solved for nin terms of /C12n: nˆn nÿ/C12n; …11:32† where /C12nis given by /C12nˆZb af…t†un…t†dt: …11:33† Finally, our solution is given by u…x†ˆf…x†‡X1 nˆ1 nun…x† n ˆf…x†‡X1 nˆ1/C12n nÿun…x†; …11:34† where /C12nis given by Eq. (11.33), and i6ˆ. When for the inhomogeneous equation is equal to one of the eigenvalues, k, of the kernel, our solution (11.31) blows up. Let us return to Eq. (11.31) and see what happens to k: kˆ/C12k‡k k=kˆ/C12k‡ k: 424SIMPLE LINEAR INTEGRAL EQUATIONS Clearly, /C12kˆ0, and kis no longer determined by /C12k. But we have, according to Eq. (11.33), Zb af…t†uk…t†dtˆ/C12kˆ0; …11:35† that is, f…x†is orthogonal to the eigenfunction uk…x†. Thus if ˆk, the inho- mogeneous equation has a solution only if f…x†is orthogonal to the correspond- ing eigenfunction uk…x†. The general solution of the equation is then u…x†ˆf…x†‡ kuk…x†‡ kX1 nˆ100Rb af…t†un…t†dt nÿkun…x†; …11:36† where the prime on the summation sign means that the term nˆkis to be omitted from the sum. In Eq. (11.36) the kremains as an undetermined constant. /C82elation bet/C119een di/C128erential and integral equations We have shown how an integral equation can be transformed into a di/C128erential equation that may be easier to solve than the original integral equation. We now show how to transform a di/C128erential equation into an integral equation. After we become familiar with the relation between di/C128erential and integral equations, we may state the physical problem in either form at will. Let us consider a linear second-order di/C128erential equation x00‡A…t†x0‡B…t†xˆ/C103…t†; …11:37† with the initial condition x…a†ˆx0;x0…a†ˆx0 0: Integrating Eq. (11.37), we obtain x0ˆÿZt aAx0dtÿZt aBxdt ‡Zt a/C103dt‡C1: The initial conditions require that C1ˆx0 0. We next integrate the first integral on the right hand side by parts and obtain x0ˆÿAxÿZt a…BÿA0†xdt‡Zt a/C103dt‡A…a†x0‡x0 0: Integrating again, we get xˆÿZt aAxdt ÿZt aZt aB…y†ÿA0…y†/C2/C3 x…y†dydt ‡Zt aZt a/C103…y†dydt‡A…a†x0‡x0 0/C2/C3 …tÿa†‡x0: 425DIFFERENTIAL AND INTEGRAL EQUATIONS Then using the relation Zt aZt af…y†dydtˆZt a…tÿy†f…y†dy; we can rewrite the last equation as x…t†ˆÿZt aA…y†‡…tÿy†B…y†ÿA0…y†/C8/C9 /C2/C3 x…y†dy ‡Zt a…tÿy†/C103…y†dy‡A…a†x0‡x0 0/C2/C3 …tÿa†‡x0; …11:38† which can be put into the form of a Volterra equation of the second kind x…t†ˆf…t†‡Zt a/C75…t;y†x…y†dy; …11:39† with /C75…t;y†ˆ… yÿt†‰B…y†ÿA0…y†Š ÿA…y†; …11:39a† f…t†ˆZt 0…tÿy†/C103…y†dy‡‰A…a†x0‡x0 0Š…tÿa†‡x0: …11:39b† /C85se of integral equations We have learned how linear integral equations of the more common types may be solved. We now show some uses of integral equations in physics; that is, we are going to state some physical problems in integral equation form. In 1823, Abel made one of the earliest applications of integral equations to a physical problem.Let us take a brief look at this old problem in mechanics. /C65bel’s integral equation Consider a particle of mass mfalling along a smooth curve in a vertical plane, the yzplane, under the influence of gravity, which acts in the negative zdirection. Conservation of energy gives 1 2m…_z2‡_y2†‡m/C103zˆ/C69; where _zˆdz=dt;and _yˆdy=dt:If the shape of the curve is given by yˆ/C70…z†,w e can write _yˆ…d/C70=dz†_z. Substituting this into the energy conservation equation and solving for _z, we obtain _zˆ 2/C69=mÿ2/C103z/C112  1‡…d/C70=dz†2q ˆ /C69=m/C103ÿz/C112 u…z†; …11:40† 426SIMPLE LINEAR INTEGRAL EQUATIONS where u…z†ˆ 1‡…d/C70=dz†2=2/C103q : If_zˆ0a n d zˆz0attˆ0, then /C69=m/C103ˆz0and Eq. (11.40) becomes _zˆz0ÿzp /C14 u…z†: Solving for time t, we obtain tˆÿZz0 zu…z†z0ÿzp dzˆZz z0u…z†z 0ÿzp dz; where zis the height the particle reaches at time t. /C67lassical simple harmonic oscillator Consider a linear oscillator /C127x‡/C332xˆ0;with x…0†ˆ0; _x…0†ˆ1: We can transform this di/C128erential equation into an integral equation. Comparing with Eq. (11.37), we have A…t†ˆ0;B…t†ˆ/C332;and /C103…t†ˆ0: Substituting these into Eq. (11.38) (or (11.39), (11.39a), and (11.39b)), we obtain the integral equation x…t†ˆt‡/C332Zt 0…yÿt†x…y†dy; which is equivalent to the original di/C128erential equation plus the initial conditions. /C81uantum simple harmonic oscillator The Schro /C200dinger equation for the energy eigenstates of the one-dimensional simple harmonic oscillator is ÿp2 2md2/C32 dx2‡1 2m/C332x2/C32ˆ/C69/C32: …11:41† Changing to the dimensionless variable yˆ m/C33=p/C112 x, Eq. (11.41) reduces to a simpler form: d2/C32 dy2‡… 2ÿy2†/C32ˆ0; …11:42† where ˆ2/C69=p/C33/C112 . Taking the Fourier transform of Eq. (11.42), we obtain d2/C103…k† dk2‡… 2ÿk2†/C103…k†ˆ0; …11:43† 427U S EO FI N T E G R A LE Q U A T I O N S where /C103…k†ˆ1  2pZ1 ÿ1/C32…y†eikydy …11:44† and we also assume that /C32and/C320vanish as y! 1 . Eq. (11.43) is formally identical to Eq. (11.42). Since quantities such as the total probability and the expectation value of the potential energy must be remain finite for finite E, we should expect /C103…k†;d/C103…k†=dk!0a sk! 1 . Thus gand/C32di/C128er at most by a normalization constant /C103…k†ˆc/C32…k†: It follows that /C32satisfies the integral equation c/C32…k†ˆ1  2pZ1 ÿ1/C32…y†eikydy: …11:45† The constant cmay be determined by substituting c/C32on the right hand side: c2/C32…k†ˆ1 2Z1 ÿ1Z1 ÿ1/C32…z†eizyeikydzdy ˆZ1 ÿ1/C32…z†…z‡k†dz ˆ/C32…ÿk†: Recall that /C32may be simultaneously chosen to be a parity eigenstate /C32…ÿx†ˆ /C32…x†. We see that eigenstates of even parity require c2ˆ1, or cˆ1; and for eigenstates of odd parity we have c2ˆÿ1, or cˆi. We shall leave the solution of Eq. (11.45), which can be approached in several ways, as an exercise for the reader. Problems 11.1 Solve the following integral equations: (a)u…x†ˆ1 2ÿx‡Z1 0u…t†dt; (b)u…x†ˆZ1 0u…t†dt; (c)u…x†ˆx‡Z1 0u…t†dt: 11.2 Solve the Fredholm equation of the second kind f…x†ˆu…x†‡Zb a/C75…x;t†u…t†dt; where f…x†ˆcosh x;/C75…x;t†ˆxt. 428SIMPLE LINEAR INTEGRAL EQUATIONS 11.3 The homogeneous Fredholm equation u…x†ˆZ=2 0sinxsintu…t†dt only has a solution for a particular value of . Find the value of and the solution corresponding to this value of . 11.4 Solve homogeneous Fredholm equation u…x†ˆR1 ÿ1…t‡x†u…t†dt. Find the values of and the corresponding solutions. 11.5 Check the convergence of the Neumann series (11.14) by the Cauchy ratio test. 11.6 Transform the following di/C128erential equations into integral equations: …a†dx dtÿxˆ0 with xˆ1 when tˆ0; …b†d2x dt2‡dx dt‡xˆ1 with xˆ0;dx dtˆ1 when tˆ0: 11.7 By using the Laplace transformation and the convolution theorem solve the equation u…x†ˆx‡Zx 0sin…xÿt†u…t†dt: 11.8 Given the Fredholm integral equation eÿx2ˆZ1 ÿ1eÿ…xÿt†2u…t†dt; apply the Fouurier convolution technique to solve it for u…t†. 11.9 Find the solution of the Fredholm equation u…x†ˆx‡Z1 0…x‡t†u…t†dt by the Schmidt–Hilbert method for not equal to an eigenvalue. Show that there are no solutions when is an eigenvalue. 429PROBLEMS 12 Elements of group theory Group theory did not find a use in physics until the advent of modern quantum mechanics in 1925. In recent years group theory has been applied to many branches of physics and physical chemistry, notably to problems of molecules, atoms and atomic nuclei. Mostly recently, group theory has been being applied in the search for a pattern of ‘family’ relationships between elementary particles. Mathematicians are generally more interested in the abstract theory of groups, but the representation theory of groups of direct use in a large variety of physical problems is more useful to physicists. In this chapter, we shall give an elementary introduction to the theory of groups, which will be needed for understanding the representation theory. /C68efinition of a group (group a/C120ioms) A group is a set of distinct elements for which a law of ‘combination’ is welldefined. Hence, before we give ‘group’ a formal definition, we must first define what kind of ‘elements’ do we mean. Any collection of objects, quantities or operators form a set, and each individual object, quantity or operator is calledan element of the set. A group is a set of elements A/C44 B/C44 /C67 ;...;finite or infinite in number, with a rule for combining any two of them to form a ‘product’, subject to the following four conditions: (1) The product of any two group elements must be a group element; that is, if AandBare members of the group, then so is the product AB. (2) The law of composition of the group elements is associative; that is, if A,B, and/C67are members of the group, then …AB†CˆA…BC†. (3) There exists a unit group element E, called the identity, such that /C69AˆA/C69ˆAfor every member of the group. 430 (4) Every element has a unique inverse, Aÿ1, such that AAÿ1ˆAÿ1Aˆ/C69. The use of the word ‘product’ in the above definition requires comment. The law of combination is commonly referred as ‘multiplication’, and so the result of a combination of elements is referred to as a ‘product’. However, the law of com- bination may be ordinary addition as in the group consisting of the set of all integers (positive, negative, and zero). Here ABˆA‡B, ‘zero’ is the identity, and Aÿ1ˆ… ÿ A†. The word ‘product’ is meant to symbolize a broad meaning of ‘multiplication’ in group theory, as will become clearer from the examples below. A group with a finite number of elements is called a finite group; and the number of elements (in a finite group) is the order of the group. A group containing an infinite number of elements is called an infinite group. An infinite group may be either discrete or continuous. If the number of theelements in an infinite group is denumerably infinite, the group is discrete; if the number of elements is non-denumerably infinite, the group is continuous. A group is called Abelian (or commutative) if for every pair of elements A,Bin the group, ABˆBA. In general, groups are not Abelian and so it is necessary to preserve carefully the order of the factors in a group ‘product’. A subgroup is any subset of the elements of a group that by themselves satisfy the group axioms with the same law of combination. Now let us consider some examples of groups. Example 12.1 The real numbers 1 and ÿ1 form a group of order two, under multiplication. The identity element is 1; and the inverse is 1 =x, where xstands for 1 or ÿ1. Example 12.2The set of all integers (positive, negative, and zero) forms a discrete infinite group under addition. The identity element is zero; the inverse of each element is its negative. The group axioms are satisfied: (1) is satisfied because the sum of any two integers (including any integer with itself) is always another integer. (2) is satisfied because the associative law of addition A‡…B‡C†ˆ …A‡B†‡Cis true for integers. (3) is satisfied because the addition of 0 to any integer does not alter it.(4) is satisfied because the addition of the inverse of an integer to the integer itself always gives 0, the identity element of our group: A‡… ÿ A†ˆ0. Obviously, the group is Abelian since A‡BˆB‡A. We denote this group by S 1. 431DEFINITION OF A GROUP (GROUP A/C88IOMS) The same set of all integers does not form a group under multiplication. Why/C63 Because the inverses of integers are not integers and so they are not members of the set. Example 12.3 The set of all rational numbers ( /C112=/C113, with /C1136ˆ0) forms a continuous infinite group under addition. It is an Abelian group, and we denote it by S2. The identity element is 0; and the inverse of a given element is its negative. Example 12.4 The set of all complex numbers …zˆx‡iy†forms an infinite group under addition. It is an Abelian group and we denote it by S3. The identity element is 0; and the inverse of a given element is its negative (that is, ÿzis the inverse ofz). The set of elements in S1is a subset of elements in S2, and the set of elements in S2is a subset of elements in S3. Furthermore, each of these sets forms a group under addition, thus S1is a subgroup of S2,a n d S2a subgroup of S3. Obviously S1is also a subgroup of S3. Example 12.5The three matrices ~Aˆ10 01 ; ~Bˆ01 ÿ1ÿ1 ; ~Cˆÿ1ÿ1 10 form an Abelian group of order three under matrix multiplication. The identity element is the unit matrix, /C69ˆ~A. The inverse of a given matrix is the inverse matrix of the given matrix: ~A ÿ1ˆ10 01 ˆ~A; ~Bÿ1ˆÿ1ÿ1 10 ˆ~C; ~Cÿ1ˆ01 ÿ1ÿ1 ˆ~B: It is straightforward to check that all the four group axioms are satisfied. We leave this to the reader. Example 12.6 The three permutation operations on three objects a;b;c ‰123Š;‰231Š;‰312Š form an Abelian group of order three with sequential performance as the law ofcombination. The operation /C911 2 3/C93 means we put the object afirst, object bsecond, and object cthird. And two elements are multiplied by performing first the operation on the 432ELEMENTS OF GROUP THEORY right, then the operation on the left. For example ‰231Љ312Šabcˆ‰231Šcabˆabc: Thus two operations performed sequentially are equivalent to the operation /C911 2 3/C93: ‰231Љ312Šˆ‰123Š: similarly ‰312Љ231Šabcˆ‰312Šbcaˆabc; that is, ‰312Љ231Šˆ‰123Š: This law of combination is commutative. What is the identity element of thisgroup/C63 And the inverse of a given element/C63 We leave the reader to answer these questions. The group illustrated by this example is known as a cyclic group of order 3, C 3. It can be shown that the set of all permutations of three objects ‰123Š;‰231Š;‰312Š;‰132Š;‰321Š;‰213Š forms a non-Abelian group of order six denoted by S3. It is called the symmetric group of three objects. Note that C3is a subgroup of S3. /C67/C121clic groups We now revisit the cyclic groups. The elements of a cyclic group can be expressed as power of a single element A, say, as A;A2;A3;...;A/C112ÿ1;A/C112ˆ/C69;pis the smal- lest integer for which A/C112ˆ/C69and is the order of the group. The inverse of Akis A/C112ÿk, that is, an element of the set. It is straightforward to check that all group axioms are satisfied. We leave this to the reader. It is obvious that cyclic groups are Abelian since AkAˆAAk…k</C112†. Example 12.7The complex numbers 1, i;ÿ1;ÿiform a cyclic group of order 3. In this case, Aˆiand/C112ˆ3:i n,nˆ0;1;2;3. These group elements may be interpreted as successive 90 8rotations in the complex plane …0; =2; ;and 3 =2†. Con- sequently, they can be represented by four 2 2 matrices. We shall come back to this later. Example 12.8 We now consider a second example of cyclic groups: the group of rotations of an equilateral triangle in its plane about an axis passing through its center that brings 433CYCLIC GROUPS it onto itself. This group contains three elements (see Fig. 12.1): /C69…ˆ08†/C58 the identity; triangle is left alone; A…ˆ120 8†/C58the triangle is rotated through 120 8counterclockwise, which sends Pto/C81,/C81to/C82, and /C82toP; B…ˆ240 8†/C58the triangle is rotated through 240 8counterclockwise, which sends Pto/C82,/C82to/C81, and /C81toP; C…ˆ360 8†/C58the triangle is rotated through 360 8counterclockwise, which sends Pback to P,/C81back to /C81and/C82back to /C82. Notice that Cˆ/C69. Thus there are only three elements represented by E,A, and B. This set forms a group of order three under addition. The reader can check that all four group axioms are satisfied. It is also obvious that operation Bis equiva- lent to performing operation Atwice (240 8ˆ120 8‡120 8†, and the operation /C67 corresponds to performing Athree times. Thus the elements of the group may be expressed as the power of the single element AasE,A,A2,A3…ˆ/C69): that is, it is a cyclic group of order three, and is generated by the element A. The cyclic group considered in Example 12.8 is a special case of groups of transformations (rotations, reflection, translations, permutations, etc.), thegroups of particular interest to physicists. A transformation that leaves a physical system invariant is called a symmetry transformation of the system. The set of all symmetry transformations of a system is a group, as illustrated by this example. /C71roup multiplication table A group of order nhasn 2products. Once the products of all ordered pairs of elements are specified the structure of a group is uniquely determined. It is some- times convenient to arrange these products in a square array called a group multi- plication table. Such a table is indicated schematically in Table 12.1. The element that appears at the intersection of the row labeled Aand the column labeled Bis the product AB, (in the table A2means AA, etc). It should be noted that all the 434ELEMENTS OF GROUP THEORY Figure 12.1. elements in each row or column of the group multiplication must be distinct: that is, each element appears once and only once in each row or column. This can be proved easily: if the same element appeared twice in a given row, the row labeledAsay, then there would be two distinct elements /C67and/C68such that ACˆAD.I f we multiply the equation by A ÿ1on the left, then we would have Aÿ1ACˆ Aÿ1AD,o r /C69Cˆ/C69D. This cannot be true unless CˆD, in contradiction to our hypothesis that /C67and/C68are distinct. Similarly, we can prove that all the elements in any column must be distinct. As a simple practice, consider the group C3of Example 12.6 and label the elements as follows ‰123Š!/C69;‰231Š!X;‰312Š!Y: If we label the columns of the table with the elements E,X,/C89and the rows with their respective inverses, E,Xÿ1,Yÿ1, the group multiplication table then takes the form shown in Table 12.2. Isomorphic groups Two groups are isomorphic to each other if the elements of one group can be put inone-to-one correspondence with the elements of the other so that the corresponding elements multiply in the same way. Thus if the elements A;B;C;...of the group /C71 435ISOMORPHIC GROUPS Table 12.1. Group multiplication table EA B /C67 ... EE A B /C67 ... AA A2AB A/C67 ... BB B A B2B/C67 ... /C67 /C67 /C67A /C67B C2... ............... Table 12.2. EX /C89 EE X /C89 Xÿ1Xÿ1E Xÿ1Y Yÿ1Yÿ1Yÿ1XE correspond respectively to the elements A0;B0;C0;...of/C710, then the equation ABˆCimplies that A0B0ˆC0, etc., and vice versa. Two isomorphic groups have the same multiplication tables except for the labels attached to the group elements. Obviously, two isomorphic groups must have the same order. Groups that are isomorphic and so have the same multiplication table are the same or identical, from an abstract point of view. That is why the concept of isomorphism is a key concept to physicists. Diverse groups of operators that act on diverse sets of objects have the same multiplication table; there is only one abstract group. This is where the value and beauty of the group theoretical method lie; the same abstract algebraic results may be applied in making predic- tions about a wide variety physical objects. The isomorphism of groups is a special instance of homomorphism, which allows many-to-one correspondence. Example 12.9 Consider the groups of Problems 12.2 and 12.4. The group /C71of Problem 12.2 consists of the four elements /C69ˆ1;Aˆi;Bˆÿ1;Cˆÿiwith ordinary multi/C45 plication as the rule of combination. The group multiplication table has the form shown in Table 12.3. The group /C710of Problem 12.4 consists of the following four elements, with matrix multiplication as the rule of combination /C690ˆ10 01 ;A0ˆ01 ÿ10 ;B0ˆÿ100ÿ1 ;C 0ˆ0ÿ1 10 : It is straightforward to check that the group multiplication table of group /C710has the form of Table 12.4. Comparing Tables 12.3 and 12.4 we can see that they have precisely the same structure. The two groups are therefore isomorphic. Example 12.10 We stated earlier that diverse groups of operators that act on diverse sets of objects have the same multiplication table; there is only one abstract group. To illustrate this, we consider, for simplicity, an abstract group of order two, /C712: that 436ELEMENTS OF GROUP THEORY Table 12.3. 1 iÿ1iE A B /C67 11 iÿ1 iE E A B /C67 ii ÿ1ÿi 1o r AA B /C67 E ÿ1 ÿ1ÿi 1 iB B /C67 E A ÿi ÿi 1 iÿ1 /C67/C67 E A B is, we make no a priori assumption about the significance of the two elements of our group. One of them must be the identity E, and we call the other X. Thus we have /C692ˆ/C69;/C69XˆX/C69ˆ/C69: Since each element appears once and only once in each row and column, the group multiplication table takes the form: We next consider some groups of operators that are isomorphic to /C712. First, consider the following two transformations of three-dimensional space into itself: (1) the transformation /C690, which leaves each point in its place, and (2) the transformation /C82, which maps the point …x;y;z†into the point …ÿx;ÿy;ÿz†. Evidently, R2ˆRR(the transformation Rfollowed by R) will bring each point back to its original position. Thus we have…/C69 0†2ˆ/C690,R/C690ˆ/C690RˆR/C690ˆR;R2ˆ/C690; and the group multiplication table has the same form as /C712: that is, the group formed by the set of the two operations /C690andRis isomorphic to /C712. We now associate with the two operations /C690andRtwo operators ^O/C690and ^OR, which act on real- or complex-valued functions of the spatial coordinates …x;y;z†, /C32…x;y;z†, with the following e/C128ects: ^O/C690/C32…x;y;z†ˆ/C32…x;y;z†; ^OR/C32…x;y;z†ˆ/C32…ÿx;ÿy;ÿz†: From these we see that …^O/C690†2ˆ ^O/C690; ^O/C690^ORˆ ^OR^O/C690ˆ ^OR; …^OR†2ˆ ^OR: 437ISOMORPHIC GROUPS Table 12.4. /C690A0B0C0 /C690/C690A0B0C0 A0A0B0C0/C690 B0B0C0/C690A0 C0C0/C690/C690B0 /C69X /C69/C69 X XX /C69 Obviously these two operators form a group that is isomorphic to /C712. These two groups (formed by the elements /C690,R, and the elements ^O/C690and ^OR, respectively) are the two representations of the abstract group /C712. These two simple examples cannot illustrate the value and beauty of the group theoretical method, but they do serve to illustrate the key concept of isomorphism. /C71roup of permutations and /C67a/C121le/C121/C39s theorem In Example 12.6 we examined briefly the group of permutations of three objects. We now come back to the general case of nobjects (1 ;2;...;n) placed in nboxes (or places) labeled 1, 2;...; n. This group, denoted by Sn, is called the sym- metric group on nobjects. It is of order n/C33 How do we know/C63 The first object may be put in any of nboxes, and the second object may then be put in any of nÿ1 boxes, and so forth: n…nÿ1†…nÿ2† 321ˆn/C33: We now define, following common practice, a permutation symbol P Pˆ123  n 1 2 3 n/C32! ; …12:1† which shifts the object in box 1 to box 1, the object in box 2 to box 2, and so forth, where 1 2 nis some arrangement of the numbers 1 ;2;3;...;n. The old notation in Example 12.6 can now be written as ‰231Šˆ123 231 : Fornobjects there are n/C33permutations or arrangements, each of which may be written in the form (12.1). Taking a specific example of three objects, we have P1ˆ123 123 ;P2ˆ123231 ;P 3ˆ123132 ; P 4ˆ123 213 ;P5ˆ123321 ;P 6ˆ123312 : For the product of two permutations P iPj…i;jˆ1;2;...;6†, we first perform the one on the right, Pj, and then the one on the left, Pi. Thus P3P6ˆ123 132123312 ˆ123213 ˆP 4: To the reader who has diculty seeing this result, let us explain. Consider the first column. We first perform P6, so that 1 is replaced by 3, we then perform P3and 3 438ELEMENTS OF GROUP THEORY is replaced by 2. So by the combined action 1 is replaced by 2 and we have the first column 1  2  : We leave the other two columns to be completed by the reader. Each element of a group has an inverse. Thus, for each permutation Pithere is Pÿ1 i, the inverse of Pi. We can use the property PiPÿ1 iˆP1to find Pÿ1 i. Let us find Pÿ1 6: Pÿ1 6ˆ312 123 ˆ123231 ˆP 2: It is straightforward to check that P6Pÿ1 6ˆP6P2ˆ123312123231 ˆ123123 ˆP 1: The reader can verify that our group S3is generated by the elements P2andP3, while P1serves as the identity. This means that the other three distinct elements can be expressed as distinct multiplicative combinations of P2andP3: P4ˆP2 2P3;P5ˆP2P3;P6ˆP22: The symmetric group Snplays an important role in the study of finite groups. Every finite group of order nis isomorphic to a subgroup of the permutation group Sn. This is known as Cayley’s theorem. For a proof of this theorem the interested reader is referred to an advanced text on group theory. In physics, these permutation groups are of considerable importance in the quantum mechanics of identical particles, where, if we interchange any two or more these particles, the resulting configuration is indistinguishable from the original one. Various quantities must be invariant under interchange or permuta- tion of the particles. Details of the consequences of this invariant property may befound in most first-year graduate textbooks on quantum mechanics that cover the application of group theory to quantum mechanics. /C83ubgroups and cosets A subset of a group /C71, which is itself a group, is called a subgroup of /C71. This idea was introduced earlier. And we also saw that C 3, a cyclic group of order 3, is a subgroup of S3, a symmetric group of order 6. We note that the order of C3is a factor of the order of S3. In fact, we will show that, in general, the order of a subgroup is a factor of the order of the full group(that is, the group from which the subgroup is derived). 439SUBGROUPS AND COSETS This can be proved as follows. Let /C71be a group of order nwith elements /C1031…ˆ/C69†, /C1032;...;/C103n:Let/C72, of order m, be a subgroup of /C71with elements /C1041…ˆ/C69†, /C1042;...;/C104m. Now form the set /C103/C104k…0km†, where gis any element of /C71not in/C72. This collection of elements is called the left-coset of /C72with respect to g(the left-coset, because gis at the left of /C104k). If such an element gdoes not exist, then Hˆ/C71, and the theorem holds trivially. Ifgdoes exist, than the elements /C103/C104kare all di/C128erent. Otherwise, we would have /C103/C104kˆ/C103/C104‘,o r /C104kˆ/C104‘, which contradicts our assumption that /C72is a group. Moreover, the elements /C103/C104kare not elements of /C72. Otherwise, /C103/C104kˆ/C104j, and we have /C103ˆ/C104j=/C104k: This implies that gis an element of /C72, which contradicts our assumption that g does not belong to /C72. This left-coset of /C72does not form a group because it does not contain the identity element ( /C1031ˆ/C1041ˆ/C69†. If it did form a group, it would require for some /C104j such that /C103/C104jˆ/C69or, equivalently, /C103ˆ/C104ÿ1 j. This requires gto be an element of /C72. Again this is contrary to assumption that gdoes not belong to /C72. Now every element gin/C71but not in /C72belongs to some coset g/C72. Thus /C71is a union of /C72and a number of non-overlapping cosets, each having mdi/C128erent elements. The order of /C71is therefore divisible by m. This proves that the order of a subgroup is a factor of the order of the full group. The ratio n/C47mis the index of/C72in/C71. It is straightforward to prove that a group of order p, where pis a prime number, has no subgroup. It could be a cyclic group generated by an element a of period p. /C67on/C106ugate classes and in/C118ariant subgroups Another way of dividing a group into subsets is to use the concept of classes. Let a,b, and ube any three elements of a group, and if bˆuÿ1au; bis said to be the transform of aby the element u;aandbare conjugate (or equivalent) to each other. It is straightforward to prove that conjugate has thefollowing three properties: (1) Every element is conjugate with itself (reflexivity). Allowing uto be the identity element E, then we have aˆ/C69 ÿ1a/C69: (2) If ais conjugate to b, then bis conjugate to a(symmetry). If aˆuÿ1bu, then bˆuauÿ1ˆ…uÿ1†ÿ1a…uÿ1†, where uÿ1is an element of /C71ifuis. 440ELEMENTS OF GROUP THEORY (3) If ais conjugate with both bandc, then bandcare conjugate with each other (transitivity). If aˆuÿ1buand bˆ/C118ÿ1c/C118, then aˆuÿ1/C118ÿ1c/C118uˆ …/C118u†ÿ1c…/C118u†, where uand/C118belong to /C71so that /C118uis also an element of /C71. We now divide our group up into subsets, such that all elements in any subset are conjugate to each other. These subsets are called classes of our group. Example 12.11 The symmetric group S3has the following six distinct elements: P1ˆ/C69;P2;P3;P4ˆP2 2P3;P5ˆP2P3;P6ˆP22; which can be separated into three conjugate classes: fP1g;fP2;P6g;fP3;P4;P5g: We now state some simple facts about classes without proofs: (a) The identity element always forms a class by itself. (b) Each element of an Abelian group forms a class by itself. (c) All elements of a class have the same period. Starting from a subgroup /C72of a group /C71, we can form a set of elements u/C104ÿ1u for each ubelong to /C71. This set of elements can be seen to be itself a group. It is a subgroup of /C71and is isomorphic to /C72. It is said to be a conjugate subgroup to /C72 in/C71. It may happen, for some subgroup /C72, that for all ubelonging to /C71, the sets /C72andu/C104uÿ1are identical. /C72is then an invariant or self-conjugate subgroup of /C71. Example 12.12 Let us revisit S3of Example 12.11, taking it as our group /C71ˆS3. Consider the subgroup HˆC3ˆfP1;P2;P6g:The following relation holds P2P1 P2 P60 BB@1 CCAPÿ1 2ˆP1P2P1 P2 P60 BB@1 CCAPÿ1 2Pÿ1 1ˆ…P1P2†P1 P2 P60 BB@1 CCA…P1P2†ÿ1 ˆP2 1P2P1 P2 P60 BB@1 CCA…P 2 1P2†ÿ1ˆP1 P6 P20 BB@1 CCA: Hence HˆC 3ˆfP1;P2;P6gis an invariant subgroup of S3. 441CONJUGATE CLASSES AND INVARIANT SUBGROUPS /C71roup representations In previous sections we have seen some examples of groups which are isomorphic with matrix groups. Physicists have found that the representation of group elements by matrices is a very powerful technique. It is beyond the scope of this text to make a full study of the representation of groups; in this section we shall make a brief study of this important subject of the matrix representations of groups. If to every element of a group /C71,/C1031;/C1032;/C1033;...;we can associate a non-singular square matrix D…/C1031†;D…/C1032†;D…/C1033†;...;in such a way that /C103i/C103jˆ/C103k implies D…/C103i†D…/C103j†ˆD…/C103k†; …12:2† then these matrices themselves form a group /C710, which is either isomorphic or homomorphic to /C71. The set of such non-singular square matrices is called a representation of group /C71. If the matrices are nn, we have an n-dimensional representation; that is, the order of the matrix is the dimension (or order) of therepresentation D n. One trivial example of such a representation is the unit matrix associated with every element of the group. As shown in Example 12.9, the four matrices of Problem 12.4 form a two-dimensional representation of the group /C71 of Problem 12.2. If there is one-to-one correspondence between each element of /C71and the matrix representation group /C710, the two groups are isomorphic, and the representation is said to be faithful (or true). If one matrix /C68represents more than one group element of /C71, the group /C71is homomorphic to the matrix representation group /C710and the representation is said to be unfaithful. Now suppose a representation of a group /C71has been found which consists of matrices DˆD…/C1031†;D…/C1032†;D…/C1033†;...;D…/C103/C112†, each matrix being of dimension n. We can form another representation D0by a similarity transformation D0…/C103†ˆSÿ1D…/C103†S; …12:3† Sbeing a non-singular matrix, then D0…/C103i†D0…/C103j†ˆSÿ1D…/C103i†SSÿ1D…/C103j†S ˆSÿ1D…/C103i†D…/C103j†S ˆSÿ1D…/C103i/C103j†S ˆD0…/C103i/C103j†: In general, representations related in this way by a similarity transformation are regarded as being equivalent. However, the forms of the individual matrices in the two equivalent representations will be quite di/C128erent. With this freedom in the 442ELEMENTS OF GROUP THEORY choice of the forms of the matrices it is important to look for some quantity that is an invariant for a given transformation. This is found in considering the traces ofthe matrices of the representation group because the trace of a matrix is invariant under a similarity transformation. It is often possible to bring, by a similarity transformation, each matrix in the representation group into a diagonal form S ÿ1DSˆD…1†0 0D…2†/C32! ; …12:4† where D…1†is of order m;m<nandD…2†is of order nÿm. Under these conditions, the original representation is said to be reducible to D…1†andD…2†. We may write this result as DˆD…1†/C8D…2†…12:5† and say that /C68has been decomposed into the two smaller representation D…1†and D…2†;/C68is often called the direct sum of D…1†andD…2†. A representation D…/C103†is called irreducible if it is not of the form (12.4) and cannot be put into this form by a similarity transformation. Irreducible representations are the simplest representations, all others may be built up from them, that is, they play the role of ‘building blocks’ for the study of group representation. In general, a given group has many representations, and it is always possible to find a unitary representation – one whose matrices are unitary. Unitary matricescan be diagonalized, and the eigenvalues can serve for the description or classi- fication of quantum states. Hence unitary representations play an especially important role in quantum mechanics. The task of finding all the irreducible representations of a group is usually very laborious. Fortunately, for most physical applications, it is sucient to know only the traces of the matrices forming the representation, for the trace of a matrix is invariant under a similarity transformation. Thus, the trace can be used to iden- tify or characterize our representation, and so it is called the character in group theory. A further simplification is provided by the fact that the character of every element in a class is identical, since elements in the same class are related to each other by a similarity transformation. If we know all the characters of one elementfrom every class of the group, we have all of the information concerning the group that is usually needed. Hence characters play an important part in the theory of group representations. However, this topic and others related to whether a given representation of a group can be reduced to one of smaller dimensions are beyond the scope of this book. There are several important theorems of representation theory, which we now state without proof. 443GROUP REPRESENTATIONS (1) A matrix that commutes with all matrices of an irreducible representation of a group is a multiple of the unit matrix (perhaps null). That is, if matrix A commutes with D…/C103†which is irreducible, D…/C103†AˆAD…/C103† for all gin our group, then Ais a multiple of the unit matrix. (2) A representation of a group is irreducible if and only if the only matrices to commute with all matrices are multiple of the unit matrix. Both theorems (1) and (2) are corollaries of Schur’s lemma. (3) Schur’s lemma: Let D…1†andD…2†be two irreducible representations of (a group /C71) dimensionality nandn0, if there exists a matrix Asuch that AD…1†…/C103†ˆD…2†…/C103†A for all /C103in the group /C71 then for n6ˆn0,Aˆ0; for nˆn0, either Aˆ0o rAis a non-singular matrix andD…1†andD…2†are equivalent representations under the similarity trans- formation generated by A. (4) Orthogonality theorem: If /C71is a group of order handD…1†andD…2†are any two inequivalent irreducible (unitary) representations, of dimensions d1and d2, respectively, then X /C103‰D…i† /C12…/C103†Š/C42D…j† /C13…/C103†ˆ/C104 d1ij /C13/C12; where D…i†…/C103†is a matrix, and D…i† /C12…/C103†is a typical matrix element. The sum runs over all gin/C71. /C83ome special groups Many physical systems possess symmetry properties that always lead to certain quantity being invariant. For example, translational symmetry (or spatial homo- geneity) leads to the conservation of linear momentum for a closed system, and rotational symmetry (or isotropy of space) leads to the conservation of angular momentum. Group theory is most appropriate for the study of symmetry. In this section we consider the geometrical symmetries. This provides more illustrations of the group concepts and leads to some special groups. Let us first review some symmetry operations. A plane of symmetry is a plane in the system such that each point on one side of the plane is the mirror image of acorresponding point on the other side. If the system takes up an identical position on rotation through a certain angle about an axis, that axis is called an axis of sym- metry. A center of inversion is a point such that the system is invariant under the operation r!ÿ r, where ris the position vector of any point in the system referred to the inversion center. If the system takes up an identical position after a rotationfollowed by an inversion, the system possesses a rotation–inversion center. 444ELEMENTS OF GROUP THEORY Some symmetry operations are equivalent. As shown in Fig. 12.2, a two-fold inversion axis is equivalent to a mirror plane perpendicular to the axis. There are two di/C128erent ways of looking at a rotation, as shown in Fig. 12.3. According to the so-called active view, the system (the body) undergoes a rotation through an angle , say, in the clockwise direction about the x3-axis. In the passive view, this is equivalent to a rotation of the coordinate system through the sameangle but in the counterclockwise sense. The relation between the new and old coordinates of any point in the body is the same in both cases: x 0 1ˆx1cos‡x2sin; x0 2ˆÿx1sin‡x2cos; x0 3ˆx3;9 >>= >>;…12:6† where the prime quantities represent the new coordinates. A general rotation, reflection, or inversion can be represented by a linear transformation of the form x0 1ˆ 11x1‡ 12x2‡ 13x3; x0 2ˆ 21x1‡ 22x2‡ 23x3; x0 3ˆ 31x1‡ 32x2‡ 33x3:9 >>= >>;…12:7† 445SOME SPECIAL GROUPS Figure 12.2. Figure 12.3. ( a) Active view of rotation; ( b) passive view of rotation. Equation (12.7) can be written in matrix form ~x0ˆ~~x …12:8† with ~ˆ111213 212223 3132330 B@1 CA; ~xˆx1 x2 x30 B@1 CA; ~x0ˆx0 1 x0 2 x0 30 B@1 CA: The matrix ~is an orthogonal matrix and the value of its determinant is 1. The ‘ÿ1’ value corresponds to an operation involving an odd number of reflections. For Eq. (12.6) the matrix ~has the form ~ˆcossin0 ÿsincos0 00 10 B@1 CA: …12:6a† For a rotation, an inversion about an axis, or a reflection in a plane through the origin, the distance of a point from the origin remains unchanged: r2ˆx2 1‡x22‡x23ˆx0 12‡x0 22‡x0 32: …12:9† /C84he symmetry group /C682;/C683 Let us now examine two simple examples of symmetry and groups. The first one is on twofold symmetry axes. Our system consists of six particles: two identicalparticles Alocated at aon the x-axis, two particles, Batbon the y-axis, and two particles /C67atcon the z-axis. These particles could be the atoms of a molecule or part of a crystal. Each axis is a twofold symmetry axis. Clearly, the identity or unit operator (no rotation) will leave the system unchanged. What rotations can be carried out that will leave our system invariant/C63 A certain com- bination of rotations of radians about the three coordinate axes will do it. The orthogonal matrices that represent rotations about the three coordinate axes canbe set up in a similar manner as was done for Eq. (12.6a), and they are ~ …†ˆ100 0ÿ10 00 ÿ10 B@1 CA; ~/C12…†ˆÿ10 0 01 0 00 ÿ10 B@1 CA; ~/C13…†ˆÿ10 0 0ÿ10 00 10 B@1 CA; where ~ is the rotational matrix about the x-axis, and ~/C12and ~/C13are the rotational matrices about y-a n d z-axes, respectively. Of course, the identity operator is a unit matrix 446ELEMENTS OF GROUP THEORY ~/C69ˆ100 010 0010 B@1 CA: These four elements form an Abelian group with the group multiplication table shown in Table 12.5. It is easy to check this group table by matrix multiplication. Or you can check it by analyzing the operations themselves, a tedious task. This demonstrates the power of mathematics: when the system becomes too complexfor a direct physical interpretation, the usefulness of mathematics shows. This symmetry group is usually labeled D 2, a dihedral group with a twofold symmetry axis. A dihedral group Dnwith an n-fold symmetry axis has naxes with an angular separation of 2 =nradians and is very useful in crystallographic study. We next consider an example of threefold symmetry axes. To this end, let us re- visit Example 12.8. Rotations of the triangle of 0 8, 120 8, 240 8, and 360 8leave the triangle invariant. Rotation of the triangle of 0 8means no rotation, the triangle is left unchanged; this is represented by a unit matrix (the identity element). Theother two orthogonal rotational matrices can be set up easily: ~AˆR z…120 8†ˆÿ1=2ÿ 3p =2 3p =2ÿ1=20 @1A; ~BˆR z…240 8†ˆÿ1=2 3p =2 ÿ3p =2ÿ1=20 @1A; and ~/C69ˆR z…0†ˆ10 01 : We notice that ~CˆRz…360 8†ˆ ~/C69. The set of the three elements …~/C69;~A;~B†forms a cyclic group C3with the group multiplication table shown in Table 12.6. The z- 447SOME SPECIAL GROUPS Table 12.5. ~/C69 ~ ~/C12 ~/C13 ~/C69 ~/C69 ~ ~/C12 ~/C13 ~ ~ ~/C69 ~/C13 ~/C12 ~/C12 ~/C12 ~/C13 ~/C69 ~ ~/C13 ~/C13 ~/C12 ~ ~/C69 axis is a threefold symmetry axis. There are three additional axes of symmetry in thexyplane: each corner and the geometric center Odefining an axis; each of these is a twofold symmetry axis (Fig. 12.4). Now let us consider reflection opera- tions. The following successive operations will bring the equilateral angle onto itself (that is, be invariant): ~/C69the identity; triangle is left unchanged; ~Atriangle is rotated through 120 8clockwise; ~Btriangle is rotated through 240 8clockwise; ~Ctriangle is reflected about axis O/C82(or the y-axis); ~Dtriangle is reflected about axis O/C81; ~/C70triangle is reflected about axis OP. Now the reflection about axis O/C82is just a rotation of 180 8about axis O/C82, thus ~CˆROR…180 8†ˆÿ10 01 : Next, we notice that reflection about axis O/C81is equivalent to a rotation of 240 8 about the z-axis followed by a reflection of the x-axis (Fig. 12.5): ~DˆRO/C81…180 8†ˆ ~C~Bˆÿ10 01ÿ1=2 3p =2 ÿ3p =2ÿ1=2/C32! ˆ1=2ÿ3p =2 ÿ3p =2ÿ1=2/C32! : 448ELEMENTS OF GROUP THEORY Table 12.6. ~/C69 ~A ~B ~/C69 ~/C69 ~A ~B ~A ~A ~B ~/C69 B ~B ~/C69 ~A Figure 12.4. Similarly, reflection about axis OPis equivalent to a rotation of 180 8followed by a reflection of the x-axis: ~/C70ˆROP…180 8†ˆ ~C~Aˆ1=2 3p =23p =2ÿ1=2/C32! : The group multiplication table is shown in Table 12.7. We have constructed a six- element non-Abelian group and a 2 2 irreducible matrix representation of it. Our group is known as D 3in crystallography, the dihedral group with a threefold axis of symmetry. One-dimensional unitrary group U…1† We now consider groups with an infinite number of elements. The group element will contain one or more parameters that vary continuously over some range so they are also known as continuous groups. In Example 12.7, we saw that the complex numbers (1 ;i;ÿ1;ÿi†form a cyclic group of order 3. These group ele- ments may be interpreted as successive 90 8rotations in the complex plane …0; =2; ;3=2†, and so they may be written as ei’with ’ˆ0,=2,,3=2. If ’is allowed to vary continuously over the range ‰0;2Š, then we will have, instead of a four-member cyclic group, a continuous group with multiplication for the composition rule. It is straightforward to check that the four group axioms are all 449SOME SPECIAL GROUPS Table 12.7. ~/C69 ~A ~B ~C ~D ~/C70 ~/C69 ~/C69 ~A ~B ~C ~D ~/C70 ~A ~A ~B ~/C69 ~D ~/C70 ~C ~B ~B ~/C69 ~A ~/C70 ~C ~D ~C ~C ~/C70 ~D ~/C69 ~B ~A ~D ~D ~C ~/C70 ~A ~/C69 ~B ~/C70 ~/C70 ~D ~C ~B ~A ~/C69 Figure 12.5. met. In quantum mechanics, ei’is a complex phase factor of a wave function, which we denote by U…’). Obviously, U…0†is an identity element. Next, U…’†U…’0†ˆei…’‡’0†ˆU…’‡’0†; andU…’‡’0†is an element of the group. There is an inverse: Uÿ1…’†ˆU…ÿ’†, since U…’†U…ÿ’†ˆU…ÿ’†U…’†ˆU…0†ˆ/C69 for any ’. The associative law is satisfied: ‰U…’1†U…’2†ŠU…’3†ˆei…’1‡’2†ei’3ˆei…’1‡’2‡’3†ˆei’1ei…’2‡’3† ˆU…’1†‰U…’2†U…’3†Š: This group is a one-dimensional unitary group; it is called U…1†. Each element is characterized by a continuous parameter ’,0’2;’can take on an infinite number of values. Moreover, the elements are di/C128erentiable: dUˆU…’‡d’†ÿU…’†ˆei…’‡d’†ÿei’ ˆei’…1‡id’†ÿei’ˆiei’d’ˆiUd’ or dU=d’ˆiU: Infinite groups whose elements are di/C128erentiable functions of their parameters are called Lie groups. The di/C128erentiability of group elements allows us to develop the concept of the generator. Furthermore, instead of studying the whole group, we can study the group elements in the neighborhood of the identity element. Thus Lie groups are of particular interest. Let us take a brief look at a few more Lie groups. Orthogonal groups SO…2†andSO…3† The rotations in an n-dimensional Euclidean space form a group, called O…n†. The group elements can be represented by nnorthogonal matrices, each with n…nÿ1†=2 independent elements (Problem 12.12). If the determinant of Ois set to be ‡1 (rotation only, no reflection), then the group is often labeled SO…n†. The label O‡ nis also often used. The elements of SO…2†are familiar; they are the rotations in a plane, say the xy plane: x0 y0/C32! ˆ~Rx y/C32! ˆcossin ÿsincos x y : 450ELEMENTS OF GROUP THEORY This group has one parameter: the angle . As we stated earlier, groups enter physics because we can carry out transformations on physical systems and the physical systems often are invariant under the transformations. Here x2‡y2is left invariant. We now introduce the concept of a generator and show that rotations of SO…2† are generated by a special 2 2 matrix ~2, where ~2ˆ0ÿi i0 : Using the Euler identity, eiˆcos‡isin, we can express the 2 2 rotation matrices R…†in exponential form: ~R…†ˆcossin ÿsincos ˆ~I2cos‡i~2sinˆei~2; where ~I2is a 2 2 unit matrix. From the exponential form we see that multi- plication is equivalent to addition of the arguments. The rotations close to theidentity element have small angles /C1290:We call ~ 2the generator of rotations for SO…2†. It has been shown that any element gof a Lie group can be written in the form /C103…1;2;...;n†ˆexpX iˆ1ii/C70i/C32! : For nparameters there are nof the quantities /C70i, and they are called the generators of the Lie group. Note that we can get ~2from the rotation matrix ~R…†by di/C128erentiation at the identity of SO…2†, that is, /C1290:This suggests that we may find the generators of other groups in a similar manner. Fornˆ3 there are three independent parameters, and the set of 3 3 ortho- gonal matrices with determinant ‡1 also forms a group, the SO…3†, its general member may be expressed in terms of the Euler angle rotation R… ; /C12; /C13 †ˆRz0…0;0; †Ry…0;/C12 ;0†Rz…0;0;/C13†; where Rzis a rotation about the z-axis by an angle /C13,Rya rotation about the y- axis by an angle /C12, and Rz0a rotation about the z0-axis (the new z-axis) by an angle . This sequence can perform a general rotation. The separate rotations can be written as ~Ry…/C12†ˆcos/C120ÿsin/C12 01 0 sin/C120 cos /C120 B@1 CA; ~Rz…/C13†ˆcos/C13 sin/C130 ÿsin/C13cos/C130 0 /C111 10 B@1 CA; 451SOME SPECIAL GROUPS ~Rx…†ˆ10 0 0c o s sin 0ÿsincos0 B@1 CA: TheSO…3†rotations leave x2‡y2‡z2invariant. The rotations Rz…/C13†form a group, called the group Rz, which is an Abelian sub- group of SO…3†. To find the generator of this group, let us take the following di/C128erentiation ÿid ~Rz…/C13†=d/C13/C13ˆ0/C12/C12ˆ0ÿi0 i00 0000 B@1 CA~Sz; where the insertion of iis to make ~SzHermitian. The rotation Rz…/C13†through an infinitesimal angle /C13can be written in terms of ~Sz: Rz…/C13†ˆ ~I3‡dRz…/C13† d/C13/C12/C12/C12/C12 /C13ˆ0/C13‡O……/C13†2†ˆ ~I3‡i/C13~Sz: A finite rotation R…/C13†may be constructed from successive infinitesimal rotations Rz…/C131‡/C132†ˆ… ~I3‡i/C131~Sz†…~I3‡i/C132~Sz†: Now let …/C13ˆ/C13=/C78for/C78rotations, with /C78!1 , then Rz…/C13†ˆ lim /C78!1~I3‡…i/C13=/C78†~Sz/C2/C3 /C78ˆexp…i~Sz†; which identifies ~Szas the generator of the rotation group Rz. Similarly, we can find the generators of the subgroups of rotations about the x-axis and the y-axis. /C84heSU…n†groups Thennunitary matrices ~Ualso form a group, the U…n†group. If there is the additional restriction that the determinant of the matrices be ‡1, we have the special unitary or unitary unimodular group, SU…n†. Each nnunitary matrix hasn2ÿ1 independent parameters (Problem 12.14). Fornˆ2w eh a v e SU…2†and possible ways to parameterize the matrix Uare ~Uˆab ÿb/C42a/C42 ; where a,bare arbitrary complex numbers and jaj2‡jbj2ˆ1. These parameters are often called the Cayley–Klein parameters, and were first introduced by Cayley and Klein in connection with problems of rotation in classical mechanics. Now let us write our unitary matrix in exponential form: ~Uˆei~H; 452ELEMENTS OF GROUP THEORY where ~His a Hermitian matrix. It is easy to show that ei~His unitary: …ei~H†‡…ei~H†ˆeÿi~H‡ei~Hˆei…~Hÿ~H‡†ˆ1: This implies that any nnunitary matrix can be written in exponential form with a particularly selected set of n2Hermitian nnmatrices, ~Hj ~Uˆexp iXn2 jˆ1j~Hj/C32! ; where the jare real parameters. The n2~Hjare the generators of the group U…n†. To specialize to SU…n†we need to meet the restriction det ~Uˆ1. To impose this restriction we need to use the identity dete~AˆeTr~A for any square matrix ~A. The proof is left as homework (Problem 12.15). Thus the condition det ~Uˆ1 requires Tr ~Hˆ0 for every ~H. Accordingly, the generators ofSU…n†are any set of nntraceless Hermitian matrices. Fornˆ2,SU…n†reduces to SU…2†, which describes rotations in two-dimen- sional complex space. The determinant is ‡1. There are three continuous para- meters (22ÿ1ˆ3). We have expressed these as Cayley–Klein parameters. The orthogonal group SO…3†, determinant ‡1, describes rotations in ordinary three- dimensional space and leaves x2‡y2‡z2invariant. There are also three inde- pendent parameters. The rotation interpretations and the equality of numbers of independent parameters suggest these two groups may be isomorphic or homo- morphic. The correspondence between these groups has been proved to be two-to-one. Thus SU…2†andSO…3†are isomorphic. It is beyond the scope of this book to reproduce the proof here. TheSU…2†group has found various applications in particle physics. For exam- ple, we can think of the proton ( p) and neutron ( n) as two states of the same particle, a nucleon /C78, and use the electric charge as a label. It is also useful to imagine a particle space, called the strong isospin space, where the nucleon statepoints in some direction, as shown in Fig. 12.6. If (or assuming that) the theory that describes nucleon interactions is invariant under rotations in strong isospin space, then we may try to put the proton and the neutron as states of a spin-likedoublet, or SU…2†doublet. Other hadrons (strong-interacting particles) can also be classified as states in SU…2†multiplets. Physicists do not have a deep under- standing of why the Standard Model (of Elementary Particles) has an SU…2† internal symmetry. Fornˆ3 there are eight independent parameters …3 2ÿ1ˆ8), and we have SU…3†, which is very useful in describing the color symmetry. 453SOME SPECIAL GROUPS /C72omogeneous /C76orent/C122 group Before we describe the homogeneous Lorentz group, we need to know the Lorentz transformation. This will bring us back to the origin of special theory of relativity. In classical mechanics, time is absolute and the Galilean trans- formation (the principle of Newtonian relativity) asserts that all inertial frames are equivalent for describing the laws of classical mechanics. But physicists in the nineteenth century found that electromagnetic theory did not seem to obey the principle of Newtonian relativity. Classical electromagnetic theory is sum- marized in Maxwell’s equations, and one of the consequences of Maxwell’sequations is that the speed of light (electromagnetic waves) is independent of the motion of the source. However, under the Galilean transformation, in a frame of reference moving uniformly with respect to the light source the light wave is no longer spherical and the speed of light is also di/C128erent. Hence, for electromagnetic phenomena, inertial frames are not equivalent and Maxwell’s equations are not invariant under Galilean transformation. A number of experiments were proposed to resolve this conflict. After the Michelson– Morley experiment failed to detect ether, physicists finally accepted that Maxwell’s equations are correct and have the same form in all inertial frames.There had to be some transformation other than the Galilean transformation that would make both electromagnetic theory and classical mechanical invar- iant. This desired new transformation is the Lorentz transformation, worked out by H. Lorentz. But it was not until 1905 that Einstein realized its full implications and took the epoch-making step involved. In his paper, ‘On the Electrodynamics of Moving Bodies’ ( The Principle of /C82elativity , Dover, New York, 1952), he developed the Special Theory of Relativity from two fundamental postulates,which are rephrased as follows: 454ELEMENTS OF GROUP THEORY Figure 12.6. The strong isospin space. (1) The laws of physics are the same in all inertial frame. No preferred inertial frame exists. (2) The speed of light in free space is the same in all inertial frames and is inde- pendent of the motion of the source (the emitting body). These postulates are often called Einstein’s principle of relativity, and they radi- cally revised our concepts of space and time. Newton’s laws of motion abolish theconcept of absolute space, because according to the laws of motion there is no absolute standard of rest. The non-existence of absolute rest means that we can- not give an event an absolute position in space. This in turn means that space is not absolute. This disturbed Newton, who insisted that there must be some abso- lute standard of rest for motion, remote stars or the ether system. Absolute space was finally abolished in its Maxwellian role as the ether. Then absolute time was abolished by Einstein’s special relativity. We can see this by sending a pulse of light from one place to another. Since the speed of light is just the distance it has traveled divided by the time it has taken, in Newtonian theory, di/C128erent observers would measure di/C128erent speeds for the light because time is absolute. Now in relativity, all observers agree on the speed of light, but they do not agree on thedistance the light has traveled. So they cannot agree on the time it has taken. That is, time is no longer absolute. We now come to the Lorentz transformation, and suggest that the reader to consult books on special relativity for its derivation. For two inertial frames with their corresponding axes parallel and the relative velocity /C118along the x 1…ˆx†axis, the Lorentz transformation has the form: x0 1ˆ/C13…x1‡i/C12x4†; x0 2ˆx2; x0 3ˆx3; x0 4ˆ/C13…x4ÿi/C12x1†; where x4ˆict;/C12ˆ/C118=c,a n d /C13ˆ1= 1ÿ/C122/C112 . We will drop the two directions perpendicular to the motion in the following discussion. For an infinitesimal relative velocity /C118, the Lorentz transformation reduces to x0 1ˆx1‡i/C12x4; x0 4ˆx4ÿi/C12x1; where /C12ˆ/C118=c;/C13ˆ1= 1ÿ…/C12†2q /C251:In matrix form we have x0 1 x0 4/C32! ˆ1 i/C12 ÿi/C12 1/C32! x1 x4/C32! : 455SOME SPECIAL GROUPS We can express the transformation matrix in exponential form: 1 i/C12 ÿi/C12 1/C32! ˆ10 01 ‡/C120i ÿi0 ˆ~I‡/C12~; where ~Iˆ10 01 ; ~ˆ0i ÿi0 : Note that ~is the negative of the Pauli spin matrix ~2. Now we have x0 1 x0 4/C32! ˆ… ~I‡/C12~†x1 x4/C32! : We can generate a finite transformation by repeating the infinitesimal transforma- tion/C78times with /C78/C12ˆ: x0 1 x0 4 ˆ ~I‡~ /C78/C78x1 x4 : In the limit as /C78!1 , lim /C78!1~I‡~ /C78/C78 ˆe~: Now we can expand the exponential in a Maclaurin series: e~ˆ~I‡~‡…~†2=2/C33‡…~†3=3/C33‡ and, noting that ~2ˆ1 and sinhˆ‡3=3/C33‡5=5/C33‡7=7/C33‡ ; cosh ˆ1‡2=2/C33‡4=4/C33‡6=6/C33‡ ; we finally obtain e~ˆ~Icosh ‡~sinh: Our finite Lorentz transformation then takes the form: x0 1 x0 2 ˆcosh isinh ÿisinhcosh  x1 x2 ; and ~is the generator of the representations of our Lorentz transformation. The transformation cosh isinh ÿisinhcosh  456ELEMENTS OF GROUP THEORY can be interpreted as the rotation matrix in the complex x4x1plane (Problem 12.16). It is straightforward to generalize the above discussion to the general case where the relative velocity is in an arbitrary direction. The transformation matrix will be a 4 4 matrix, instead of a 2 2 matrix one. For this general case, we have to take x2- and x3-axes into consideration. Problems 12.1. Show that (a) the unit element (the identity) in a group is unique, and (b) the inverse of each group element is unique. 12.2. Show that the set of complex numbers 1 ;i;ÿ1, and ÿiform a group of order four under multiplication. 12.3. Show that the set of all rational numbers, the set of all real numbers, and the set of all complex numbers form infinite Abelian groups underaddition. 12.4. Show that the four matrices ~Aˆ10 01 ; ~Bˆ01 ÿ10 ; ~Cˆÿ100ÿ1 ; ~Dˆ0ÿ1 10 form an Abelian group of order four under multiplication. 12.5. Show that the set of all permutations of three objects ‰123Š;‰231Š;‰312Š;‰132Š;‰321Š;‰213Š forms a non-Abelian group of order six, with sequential performance as the law of combination. 12.6. Given two elements AandBsubject to the relations A 2ˆB2ˆ/C69(the identity), show that:(a)AB6ˆBA, and (b) the set of six elements /C69;A;B;A 2;AB;BAform a group. 12.7. Show that the set of elements 1 ;A;A2;...;Anÿ1,Anˆ1, where Aˆe2i=n forms a cyclic group of order nunder multiplication. 12.8. Consider the rotations of a line about the z-axis through the angles =2; ;3=2;and 2 in the xyplane. This is a finite set of four elements, the four operations of rotating through =2; ;3=2, and 2 . Show that this set of elements forms a group of order four under addition. 12.9. Construct the group multiplication table for the group of Problem 12.2. 12.10. Consider the possible rearrangement of two objects. The operation /C69/C112 leaves each object in its place, and the operation I/C112interchanges the two objects. Show that the two operations form a group that is isomorphic to /C712. 457PROBLEMS Next, we associate with the two operations two operators ^O/C69/C112and ^OI/C112, which act on the real or complex function f…x1;y1;z1;x2;y2;z2†with the following e/C128ects: ^O/C69/C112fˆf; ^OI/C112f…x1;y1;z1;x2;y2;z2†ˆf…x2;y2;z2;x1;y1;z1†: Show that the two operators form a group that is isomorphic to /C712. 12.11. Verify that the multiplication table of S3has the form: 12.12. Show that an nnorthogonal matrix has n…nÿ1†=2 independent elements. 12.13. Show that the 2 2 matrix 2can be obtained from the rotation matrix R…†by di/C128erentiation at the identity of SO…2†, that is, ˆ0. 12.14. Show that an nnunitary matrix has n2ÿ1 independent parameters. 12.15. Show that det e~AˆeTr ~Awhere ~Ais any square matrix. 12.16. Show that the Lorentz transformation x0 1ˆ/C13…x1‡i/C12x4†; x0 2ˆx2; x0 3ˆx3; x0 4ˆ/C13…x4ÿi/C12x1† corresponds to an imaginary rotation in the x4x1plane. (A detailed dis- cussion of this can be found in the book /C67lassical Mechanics , by Tai L. Chow, John Wiley, 1995.) 458ELEMENTS OF GROUP THEORY P1P2P3P4P5P6 P1 P1P2P3P4P5P6 P2 P2P1P6P5P6P4 P3 P3P4P5P6P2P1 P4 P4P5P3P1P6P2 P5 P5P3P4P2P1P6 P6 P6P2P1P3P4P5 13 /C78umerical methods Very few of the mathematical problems which arise in physical sciences and engineering can be solved analytically. Therefore, a simple, perhaps crude, tech- nique giving the desired values within specified limits of tolerance is often to be preferred. We do not give a full coverage of numerical analysis in this chapter; but some methods for numerically carrying out the processes of interpolation, finding roots of equations, integration, and solving ordinary di/C128erential equations will be presented. Interpolation In the eighteenth century Euler was probably the first person to use the interpola- tion technique to construct planetary elliptical orbits from a set of observed positions of the planets. We discuss here one of the most common interpolation techniques: the polynomial interpolation. Suppose we have a set of observed or measured data …x0;y0†,…x1;y1†;...;…xn;yn†, how do we represent them by a smooth curve of the form yˆf…x†/C63 For analytical convenience, this smooth curve is usually assumed to be polynomial: f…x†ˆa0‡a1x1‡a2x2‡‡ anxn…13:1† and we use the given points to evaluate the coecients a0;a1;...;an: f…x0†ˆa0‡a1x0‡a2x2 0‡‡ anxn0ˆy0; f…x1†ˆa0‡a1x1‡a2x2 1‡‡ anxn1ˆy1; ... f…xn†ˆa0‡a1xn‡a2x2 n‡‡ anxnnˆyn:9 >>>>>= >>>>>;…13:2† This provides n‡1 equations to solve for the n‡1 coecients a 0;a1;...;an: However, straightforward evaluation of coecients in the way outlined above 459 is rather tedious, as shown in Problem 13.1, hence many shortcuts have been devised, though we will not discuss these here because of limited space. Finding roots of equations A solution of an equation f…x†ˆ0 is sometimes called a root, where f…x†is a real continuous function. If f…x†is suciently complicated that a direct solution may not be possible, we can seek approximate solutions. In this section we will sketch some simple methods for determining the approximate solutions of algebraic and transcendental equations. A polynomial equation is an algebraic equation. An equation that is not reducible to an algebraic equation is called transcendental. Thus, tan xÿxˆ0a n d ex‡2c o s xˆ0 are transcendental equations. Graphical methods The approximate solution of the equation f…x†ˆ0 …13:3† can be found by graphing the function yˆf…x†and reading from the graph the values of xfor which yˆ0. The graphing procedure can often be simplified by first rewriting Eq. (13.3) in the form /C103…x†ˆ/C104…x†… 13:4† and then graphing yˆ/C103…x†andyˆ/C104…x†. The xvalues of the intersection points of the two curves gives the approximate values of the roots of Eq. (13.4). As anexample, consider the equation f…x†ˆx 3ÿ146:25xÿ682:5ˆ0; we can graph yˆx3ÿ146:25xÿ682:5 to find its roots. But it is simpler to graph the two curves yˆx3…a cubic † and yˆ146:25x‡682:5…a straight line †: See Fig. 13.1. There is one drawback of graphical methods: that is, they require plotting curves on a large scale to obtain a high degree of accuracy. To avoid this, methods of successive approximations (or simple iterative methods) have been devised, and we shall sketch a couple of these in the following sections. 460NUMERICAL METHODS Method of linear interpolation (method of false position) Make an initial guess of the root of Eq. (13.3), say x0, located between x1andx2, and in the interval ( x1;x2) the graph of yˆf…x†has the appearance as shown in Fig. 13.2. The straight line connecting P1andP2cuts the x-axis at point x3;which is usually closer to x0than either x1orx2. From similar triangles x3ÿx1 ÿf…x1†ˆx2ÿx1 f…x2†; and solving for x3we get x3ˆx1f…x2†ÿx2f…x1† f…x2†ÿf…x1†: Now the straight line connecting the points P3andP2intersects the x-axis at point x4, which is a closer approximation to x0than x3. By repeating this process we obtain a sequence of values x3;x4;...;xnthat generally converges to the root of the equation. The iterative method described above can be simplified if we rewrite Eq. (13.3) in the form of Eq. (13.4). If the roots of /C103…x†ˆc …13:5† can be determined for every real c, then we can start the iterative process as follows. Let x1be an approximate value of the root x0of Eq. (13.3) (and, of 461METHOD OF LINEAR INTERPOLATION Figure 13.1. course, also equation 13.4). Now setting xˆx1on the right hand side of Eq. (13.4) we obtain the equation /C103…x†ˆ/C104…x1†; …13:6† which by hypothesis we can solve. If the solution is x2, we set xˆx2on the right hand side of Eq. (13.4) and obtain /C103…x†ˆ/C104…x2†: …13:7† By repeating this process, we obtain the nth approximation /C103…x†ˆ/C104…xnÿ1†: …13:8† From geometric considerations or interpretation of this procedure, we can see that the sequence x1;x2;...;xnconverges to the root xˆ0 if, in the interval 2jx1ÿx0jcentered at x0, the following conditions are met: …1†j/C1030…x†j/C62j/C1040…x†j;and …2†The derivatives are bounded :) …13:9† Example 13.1 Find the approximate values of the real roots of the transcendental equation exÿ4xˆ0: Solution: Letg…x†ˆxand h…x†ˆex=4;so the original equation can be rewrit- ten as xˆex=4: 462NUMERICAL METHODS Figure 13.2. According to Eq. (13.8) we have xn‡1ˆexn=4;nˆ1;2;3;...: …13:10† There are two roots (see Fig. 13.3), with one around xˆ0:3:If we take it as x1, then we have, from Eq. (13.10) x2ˆex1=4ˆ0:3374 ; x3ˆex2=4ˆ0:3503 ; x4ˆex3=4ˆ0:3540 ; x5ˆex4=4ˆ0:3565 ; x6ˆex5=4ˆ0:3571 ; x7ˆex6=4ˆ0:3573 : The computations can be terminated at this point if only three-decimal-place accuracy is required. The second root lies between 2 and 3. As the slope of yˆ4xis less than that of yˆex, the first condition of Eq. (13.9) cannot be met, so we rewrite the original equation in the form exˆ4x;orxˆlog 4x 463Figure 13.3.METHOD OF LINEAR INTERPOLATION and take /C103…x†ˆx,/C104…x†ˆlog 4 x. We now have xn‡1ˆlog 4 xn;nˆ1;2;...: If we take x1ˆ2:1, then x2ˆlog 4 x1ˆ2:12823 ; x3ˆlog 4 x2ˆ2:14158 ; x4ˆlog 4 x3ˆ2:14783 ; x5ˆlog 4 x4ˆ2:15075 ; x6ˆlog 4 x5ˆ2:15211 ; x7ˆlog 4 x6ˆ2:15303 ; x8ˆlog 4 x7ˆ2:15316 ; and we see that the value of the root correct to three decimal places is 2.153. Ne/C119ton’s method In Newton’s method, the successive terms in the sequence of approximate values x1;x2;...;xnthat converges to the root is obtained by the intersection with the x- axis of the tangent line to the curve yˆf…x†. Fig. 13.4 shows a portion of the graph of f…x†close to one of its roots, x0. We start with x1, an initial guess of the value of the root x0. Now the equation of the tangent line to yˆf…x†atP1is yÿf…x1†f0…x1†…xÿx1†: …13:11† This tangent line intersects the x-axis at x2that is a better approximation to the root than x1. To find x2, we set yˆ0 in Eq. (13.11) and find x2ˆx1ÿf…x1†=f0…x1† 464NUMERICAL METHODS Figure 13.4. provided f0…x1†6 ˆ0. The equation of the tangent line at P2is yÿf…x2†ˆf0…x2†…xÿx2† and it intersects the x-axis at x3: x3ˆx2ÿf…x2†=f0…x2†: This process is continued until we reach the desired level of accuracy. Thus, in general xn‡1ˆxnÿf…xn† f0…xn†;nˆ1;2;...: …13:12† Newton’s method may fail if the function has a point of inflection, or other bad behavior, near the root. To illustrate Newton’s method, let us consider the follow-ing trivial example. Example 13.2 Solve, by Newton’s method, x 3ÿ2ˆ0. Solution: Here we have yˆx3ÿ2. If we take x1ˆ1:5 (note that 1 <21=3<3†, then Eq. (13.12) gives x2ˆ1:296296296 ; x3ˆ1:260932225 ; x4ˆ1:259921861 ; x5ˆ1:25992105 x6ˆ1:25992105) repetition : Thus, to eight-decimal-place accuracy, 21=3ˆ1:25992105. When applying Newton’s method, it is often convenient to replace f0…xn†by f…xn‡†ÿf…xn† ; with small. Usually ˆ0:001 will give good accuracy. Eq. (13.12) then reads xn‡1ˆxnÿf…xn† f…xn‡†ÿf…xn†;nˆ1;2;...: …13:13† Example 13.3Solve the equation x 2ÿ2ˆ0. 465METHOD OF LINEAR INTERPOLATION Solution: Here f…x†ˆx2ÿ2. Take x1ˆ1a n d ˆ0:001, then Eq. (13.13) gives x2ˆ1:499750125 ; x3ˆ1:416680519 ; x4ˆ1:414216580 ; x5ˆ1:414213563 ; x6ˆ1:414113562 x7ˆ1:414113562) x6ˆx7: Numerical integration Very often definite integrations cannot be done in closed form. When this happens we need some simple and useful techniques for approximating definite integrals. In this section we discuss three such simple and useful methods. /C84he rectangular rule The reader is familiar with the interpretation of the definite integralRb af…x†dxas the area under the curve yˆf…x†between the limits xˆaandxˆb: Zb af…x†dxˆXn iˆ1f… i†…xiÿxiÿ1†; where xiÿ1 ixi;aˆx0<x1<x2<<xnˆb:We can obtain a good approximation to this definite integral by simply evaluating such an area underthe curve yˆf…x†. We can divide the interval axbintonsubintervals of length /C104ˆ…bÿa†=n, and in each subinterval, the function f… i†is replaced by a 466NUMERICAL METHODS Figure 13.5. straight line connecting the values at each head or end of the subinterval (or at the center point of the interval), as shown in Fig. 13.5. If we choose the head, iˆxiÿ1, then we have Zb af…x†dx/C25/C104…y0‡y1‡‡ ynÿ1†; …13:14† where y0ˆf…x0†;y1ˆf…x1†;...;ynÿ1ˆf…xnÿ1†. This method is called the rec- tangular rule. It will be shown later that the error decreases as n2. Thus, as nincreases, the error decreases rapidly. /C84he trape/C122oidal rule The trapezoidal rule evaluates the small area of a subinterval slightly di/C128erently.The area of a trapezoid as shown in Fig. 13.6 is given by 1 2/C104…Y1‡Y2†: Thus, applied to Fig. 13.5, we have the approximation Zb af…x†dx/C25…bÿa† n…1 2y0‡y1‡y2‡‡ ynÿ1‡12yn†: …13:15† What are the upper and lower limits on the error of this method/C63 Let us first calculate the error for a single subinterval of length /C104…ˆ …bÿa†=n†. Writing xi‡/C104ˆzand/C34i…z†for the error, we have Zz xif…x†dxˆ/C104 2‰yi‡yzЇ/C34i…z†; 467NUMERICAL INTEGRATION Figure 13.6. where yiˆf…xi†;yzˆf…z†.O r /C34i…z†ˆZz xif…x†dxÿ/C104 2‰f…xi†ÿf…z†Š ˆZz xif…x†dxÿzÿxi 2‰f…xi†ÿf…z†Š: Di/C128erentiating with respect to z: /C340 i…z†ˆf…z†ÿ‰f…xi†‡f…z†Š=2ÿ…zÿxi†f0…z†=2: Di/C128erentiating once again, /C3400 i…z†ˆÿ … zÿxi†f00…z†=2: IfmiandMiare, respectively, the minimum and the maximum values of f00…z†in the subinterval /C91 xi;z/C93, we can write zÿxi 2miÿ/C3400 i…z†zÿxi 2Mi: Anti-di/C128erentiation gives …zÿxi†2 4miÿ/C340 i…z†…zÿxi†2 4Mi Anti-di/C128erentiation once more gives …zÿxi†3 12miÿ/C34i…z†…zÿxi†3 12Mi: or, since zÿxiˆ/C104, /C1043 12miÿ/C34i/C1043 12Mi: Ifmand Mare, respectively, the minimum and the maximum of f00…z†in the interval /C91 a;b/C93 then /C1043 12mÿ/C34i/C1043 12M for all i: Adding the errors for all subintervals, we obtain /C1043 12nmÿ/C34/C1043 12nM or, since /C104ˆ…bÿa†=n; …bÿa†3 12n2mÿ/C34…bÿa†3 12n2M: …13:16† 468NUMERICAL METHODS Thus, the error decreases rapidly as nincreases, at least for twice-di/C128erentiable functions. Simpson’s rule Simpson’s rule provides a more accurate and useful formula for approximating a definite integral. The interval axbis subdivided into an even number of subintervals. A parabola is fitted to points a,a‡/C104,a‡2/C104; another to a‡2/C104, a‡3/C104,a‡4/C104; and so on. The area under a parabola, as shown in Fig. 13.7, is (Problem 13.8) /C104 3…y1‡4y2‡y3†: Thus, applied to Fig. 13.5, we have the approximation Zb af…x†dx/C25/C1043…y 0‡4y1‡2y2‡4y3‡2y4‡‡ 2ynÿ2‡4ynÿ1‡yn†;…13:17† with neven and /C104ˆ…bÿa†=n. The analysis of errors for Simpson’s rule is fairly involved. It has been shown that the error is proportional to /C1044(or inversely proportional to n4). There are other methods of approximating integrals, but they are not so simple as the above three. The method called Gaussian quadrature is very fast but more involved to implement. Many textbooks on numerical analysis cover this method. Numerical solutions of di/C128erential equations We noted in Chapter 2 that the methods available for the exact solution of di/C128er-ential equations apply only to a few, principally linear, types of di/C128erential equa- tions. Many equations which arise in physical science and in engineering are not solvable by such methods and we are therefore forced to find ways of obtaining approximate solutions of these di/C128erential equations. The basic idea of approxi- mate solutions is to specify a small increment hand to obtain approximate values of a solution yˆy…x†atx 0,x0‡/C104,x0‡2/C104;...: 469NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS Figure 13.7. The first-order ordinary di/C128erential equation dy dxˆf…x;y†; …13:18† with the initial condition yˆy0when xˆx0, has the solution yÿy0ˆZx x0f…t;y…t††dt: …13:19† This integral equation cannot be evaluated because the value of yunder the integral sign is unknown. We now consider three simple methods of obtaining approximate solutions: Euler’s method, Taylor series method, and the Runge– Kutta method. Euler’s method Euler proposed the following crude approach to finding the approximate solution.He began at the initial point …x 0;y0†and extended the solution to the right to the point x1ˆx0‡/C104, where /C104is a small quantity. In order to use Eq. (13.19) to obtain the approximation to y…x1†, he had to choose an approximation to fon the interval ‰x0;x1Š. The simplest of all approximations is to use f…t;y…t†† ˆ f…x0;y0†. With this choice, Eq. (13.19) gives y…x1†ˆy0‡Zx1 x0f…x0;y0†dtˆy0‡f…x0;y0†…x1ÿx0†: Letting y1ˆy…x1†, we have y1ˆy0‡f…x0;y0†…x1ÿx0†: …13:20† From y1;y0 1ˆf…x1;y1†can be computed. To extend the approximate solution further to the right to the point x2ˆx1‡/C104, we use the approximation: f…t;y…t†† ˆy0 1ˆf…x1;y1†. Then we obtain y2ˆy…x2†ˆy1‡Zx2 x1f…x1;y1†dtˆy1‡f…x1;y1†…x2ÿx1†: Continuing in this way, we approximate y3,y4, and so on. There is a simple geometrical interpretation of Euler’s method. We first note that f…x0;y0†ˆy0…x0†, and that the equation of the tangent line at the point …x0;y0†to the actual solution curve (or the integral curve) yˆy…x†is yÿy0ˆZx x0f…t;y…t††dtˆf…x0;y0†…xÿx0†: Comparing this with Eq. (13.20), we see that …x1;y1†lies on the tangent line to the actual solution curve at ( x0;y0). Thus, to move from point ( x0;y0) to point ( x1;y1) we proceed along this tangent line. Similarly, to move to point ( x2;y2†we proceed parallel to the tangent line to the solution curve at ( x1;y1†, as shown in Fig. 13.8. 470NUMERICAL METHODS The merit of Euler’s method is its simplicity, but the successive use of the tangent line at the approximate values y1;y2;...can accumulate errors. The accu- racy of the approximate vale can be quite poor, as shown by the following simple example. Example 13.4 Use Euler’s method to approximate solution to y0ˆx2‡y;y…1†ˆ3 on interval ‰1;2Š: Solution: Using hˆ0:1, we obtain Table 13.1. Note that the use of a smaller step-size hwill improve the accuracy. Euler’s method can be improved upon by taking the gradient of the integral curve as the means of obtaining the slopes at x0andx0‡/C104, that is, by using the 471NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS Table 13.1. xy (Euler) y(actual) 1.0 3 3 1.1 3.4 3.431371.2 3.861 3.931221.3 4.3911 4.50887 1.4 4.99921 5.1745 1.5 5.69513 5.939771.6 6.48964 6.81695 1.7 7.39461 7.82002 1.8 8.42307 8.964331.9 9.58938 10.2668 2.0 10.9093 11.7463 Figure 13.8. approximate value obtained for y1, we obtain an improved value, denoted by …y1†1: …y1†1ˆy0‡1 2ff…x0;y0†‡f…x0‡/C104;y1†g: …13:21† This process can be repeated until there is agreement to a required degree of accuracy between successive approximations. /C84he three-term /C84aylor series method The rationale for this method lies in the three-term Taylor expansion. Let ybe the solution of the first-order ordinary equation (13.18) for the initial condition yˆy0when xˆx0and suppose that it can be expanded as a Taylor series in the neighborhood of x0.I fyˆy1when xˆx0‡/C104, then, for suciently small values of /C104,w eh a v e y1ˆy0‡/C104dy dx 0‡/C1042 2/C33d2y dx2/C32! 0‡/C1043 3/C33d3y dx3/C32! 0‡ : …13:22† Now dy dxˆf…x;y†; d2y dx2ˆ/C64f /C64x‡dy dx/C64f /C64yˆ/C64f /C64x‡f/C64f /C64y; and d3y dx3ˆ/C64 /C64x‡f/C64 /C64y/C64f /C64x‡f/C64f /C64y ˆ/C642f /C64x2‡/C64f /C64y/C64f /C64y‡2f/C642f /C64x/C64y‡f/C64f /C64y2 ‡f2/C642f /C64y2: Equation (13.22) can be rewritten as y1ˆy0‡/C104f…x0;y0†‡/C1042 2/C64f…x0;y0† /C64x‡f…x0;y0†/C64f…x0;y0† /C64y ; where we have dropped the /C1043term. We now use this equation as an iterative equation: yn‡1ˆyn‡/C104f…xn;yn†‡/C1042 2/C64f…xn;yn† /C64x‡f…xn;yn†/C64f…xn;yn† /C64y : …13:23† That is, we compute y1ˆy…x0‡/C104†from y0,y2ˆy…x1‡/C104†from y1by replacing x byx1, and so on. The error in this method is proportional /C1043. A good approxima- 472NUMERICAL METHODS tion can be obtained for ynby summing a number of terms of the Taylor’s expansion. To illustrate this method, let us consider a very simple example. Example 13.5 Find the approximate values of y1through y10for the di/C128erential equation y0ˆx‡y, with the initial condition x0ˆ1:0 and yˆÿ2:0. Solution: Now f…x;y†ˆx‡y;/C64f=/C64xˆ/C64f=/C64yˆ1 and Eq. (13.23) reduces to yn‡1ˆyn‡/C104…xn‡yn†‡/C1042 2…1‡xn‡yn†: Using this simple formula with /C104ˆ0:1 we obtain the results shown in Table 13.2. /C84he Runge/C177/C75utta method In practice, the Taylor series converges slowly and the accuracy involved is not very high. Thus we often resort to other methods of solution such as the Runge– Kutta method, which replaces the Taylor series, Eq. (13.23), with the following formula: yn‡1ˆyn‡/C104 6…k1‡4k2‡k3†; …13:24† where k1ˆf…xn;yn†; …13:24a† k2ˆf…xn‡/C104=2;yn‡/C104k1=2†; …13:24b† k3ˆf…xn‡/C104;y0‡2/C104k2ÿ/C104k1†: …13:24c† This approximation is equivalent to Simpson’s rule for the approximate integration of f…x;y†, and it has an error proportional to /C1044. A beauty of the 473NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS Table 13.2. n xn yn yn‡1 0 1.0 ÿ2.0 ÿ2.1 1 1.1 ÿ2.1 ÿ2.2 2 1.2 ÿ2.2 ÿ2.3 3 1.3 ÿ2.3 ÿ2.4 4 1.4 ÿ2.4 ÿ2.5 5 1.5 ÿ2.5 ÿ2.6 Runge–Kutta method is that we do not need to compute partial derivatives, but it becomes rather complicated if pursued for more than two or three steps. The accuracy of the Runge–Kutta method can be improved with the following formula: yn‡1ˆyn‡/C104 6…k1‡2k2‡2k3‡k4†; …13:25† where k1ˆf…xn;yn†; …13:25a† k2ˆf…xn‡/C104=2;yn‡/C104k1=2†; …13:25b† k3ˆf…xn‡/C104;y0‡/C104k2=2†; …13:25c† k4ˆf…xn‡/C104;yn‡/C104k3†: …13:25d† With this formula the error in yn‡1is of order /C1045. You may wonder how these formulas are established. To this end, let us go back to Eq. (13.22), the three-term Taylor series, and rewrite it in the form y1ˆy0‡/C104f0‡…1=2†/C1042…A0‡f0B0†‡…1=6†/C1043…C0‡2f0D0‡f2 0/C690 ‡A0B0‡f0B2 0†‡O…/C1044†; …13:26† where Aˆ/C64f /C64x;Bˆ/C64f /C64y;Cˆ/C642f /C64x2;Dˆ/C642f /C64x/C64y;/C69ˆ/C642f /C64y2 and the subscript 0 denotes the values of these quantities at …x0;y0†. Now let us expand k1;k2,a n d k3in the Runge–Kutta formula (13.24) in powers ofhin a similar manner: k1ˆ/C104f…x0;y0†; k2ˆf…x0‡/C104=2;y0‡k1/C104=2†; ˆf0‡1 2/C104…A0‡f0B0†‡18/C104 2…C0‡2f0D0‡f2 0/C690†‡O…/C1043†: Thus 2k2ÿk1ˆf0‡/C104…A0‡f0B0†‡ and d d/C104…2k2ÿk1† /C104ˆ0ˆf0;d2 d/C1042…2k2ÿk1†/C32! /C104ˆ0ˆ2…A0‡f0B0†: 474NUMERICAL METHODS Then k3ˆf…x0‡/C104;y0‡2/C104k2ÿ/C104k1† ˆf0‡/C104…A0‡f0B0†‡…1=2†/C1042fC0‡2f0D0‡f2 0/C690‡2B0…A0‡f0B0†g ‡O…/C1043†: and …1=6†…k1‡4k2‡k3†ˆ/C104f0‡…1=2†/C1042…A0‡f0B0† ‡…1=6†/C1043…C0‡2f0D0‡f2 0/C690‡A0B0‡f0B2 0†‡O…/C1044†: Comparing this with Eq. (13.26), we see that it agrees with the Taylor series expansion (up to the term in /C1043) and the formula is established. Formula (13.25) can be established in a similar manner by taking one more term of the Taylor series. Example 13.6 Using the Runge–Kutta method and /C104ˆ0:1, solve y0ˆxÿy2=10;x0ˆ0;y0ˆ1: Solution: With hˆ0:1,h4ˆ0:0001 and we may use the Runge–Kutta third- order approximation. First step: x0ˆ0;y0ˆ0;f0ˆÿ0:1; k1ˆÿ0:1,y0‡/C104k1=2ˆ0:995; k2ˆÿ0:049;2k2ÿk1ˆ0:002;k3ˆ0; y1ˆy0‡/C104 6…k1‡4k2‡k1†ˆ0:9951 : Second step: x1ˆx0‡/C104ˆ0:1,y1ˆ0:9951, f1ˆ0:001, k1ˆ0:001,y1‡/C104k1=2ˆ0:9952 ; k2ˆ0:051, 2 k2ÿk1ˆ0:101,k3ˆ0:099, y2ˆy1‡/C1046…k 1‡4k2‡k1†ˆ1:0002 : Third step: x2ˆx1‡/C104ˆ0:2;y2ˆ1:0002, f2ˆ0:1, k1ˆ0:1,y2‡/C104k1=2ˆ1:0052, k2ˆ0:149;2k2ÿk1ˆ0:198;k3ˆ0:196; y3ˆy2‡/C1046…k 1‡4k2‡k1†ˆ1:0151 : 475NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS Equations of higher order/C46 System of equations The methods in the previous sections can be extended to obtain numerical solu- tions of equations of higher order. An nth-order di/C128erential equation is equivalent tonfirst-order di/C128erential equations in n‡1 variables. Thus, for instance, the second-order equation y00ˆf…x;y;y0†; …13:27† with initial conditions y…x0†ˆy0;y0…x0†ˆy0 0; …13:28† can be written as a system of two equations of first order by setting y0ˆu; …13:29† then Eqs. (13.27) and (13.28) become u0ˆf…x;y;u†; …13:30† y…x0†ˆy0;u…x0†ˆu0: …13:31† The two first-order equations (13.29) and (13.30) with the initial conditions (13.31) are completely equivalent to the original second-order equation (13.27) with the initial conditions (13.28). And the methods in the previous sections for determining approximate solutions can be extended to solve this system of two first-order equations. For example, the equation y00ÿyˆ2; with initial conditions y…0†ˆÿ 1;y0…0†ˆ1; is equivalent to the system y0ˆx‡u;u0ˆ1‡y; with y…0†ˆÿ 1;u…0†ˆ1: These two first-order equations can be solved with Taylor’s method (Problem 13.12). The simple methods outlined above all have the disadvantage that the error in approximating to values of yis to a certain extent cumulative and may become large unless some form of checking process is included. For this reason, methodsof solution involving finite di/C128erence are devised, most of them being variations of the Adams–Bashforth method that contains a self-checking process. This method 476NUMERICAL METHODS is quite involved and because of limited space we shall not cover it here, but it is discussed in any standard textbook on numerical analysis. Least-squares fit We now look at the problem of fitting of experimental data. In some experimentalsituations there may be underlying theory that suggests the kind of function to be used in fitting the data. Often there may be no theory on which to rely in selecting a function to represent the data. In such circumstances a polynomial is often used.We saw earlier that the m‡1 coecients in the polynomial yˆa 0‡a1x‡‡ amxm can always be determined so that a given set of m‡1 points ( xi;yi), where the xs may be unequal, lies on the curve described by the polynomial. However, whenthe number of points is large, the degree mof the polynomial is high, and an attempt to fit the data by using a polynomial is very laborious. Furthermore, theexperimental data may contain experimental errors, and so it may be more sensible to represent the data approximately by some function yˆf…x†that contains a few unknown parameters. These parameters can then be determinedso that the curve yˆf…x†fits the data. How do we determine these unknown parameters/C63 Let us represent a set of experimental data ( x i;yi), where iˆ1;2;...;n, by some function yˆf…x†that contains rparameters a1;a2;...;ar. We then take the deviations (or residuals) diˆf…xi†ÿyi …13:32† and form the weighted sum of squares of the deviations SˆXn iˆ1/C119i…di†2ˆXn iˆ1/C119i‰f…xi†ÿyiŠ2; …13:33† where the weights /C119iexpress our confidence in the accuracy of the experimental data. If the points are equally weighted, the /C119s can all be set to 1. It is clear that the quantity Sis a function of as:SˆS…a1;a2;...;ar†:We can now determine these parameters so that Sis a minimum: /C64S /C64a1ˆ0;/C64S /C64a2ˆ0;...;/C64S /C64arˆ0: …13:34† The set of requations (13.34) is called the normal equations and serves to determine the runknown asi nyˆf…x†. This particular method of determining the unknown as is known as the method of least squares. 477LEAST-SQUARES FIT We now illustrate the construction of the normal equations with the simplest case: yˆf…x†is a linear function: yˆa1‡a2x: …13:35† The deviations diare given by diˆ…a1‡a2x†ÿyi and so, assuming /C119iˆ1 SˆXn iˆ1d2 iˆ…a1‡a2x1ÿy1†2‡…a1‡a2x2ÿy2†2‡‡… a1‡a2xrÿyr†2: We now find the partial derivatives of S with respect to a1anda2and set these to zero: /C64S=/C64a1ˆ2…a1‡a2x1ÿy1†‡2…a1‡a2x2ÿy2†‡‡ 2…a1‡a2xnÿyn†ˆ0; /C64S=/C64a2ˆ2x1…a1‡a2x1ÿy1†‡2x2…a1‡a2x2ÿy2†‡‡ 2xn…a1‡a2xnÿyn† ˆ0: Dividing out the factor 2 and collecting the coecients of a1anda2, we obtain na1‡Xn iˆ1xi/C32! a2ˆXn iˆ1y1; …13:36† Xn iˆ1xi/C32! a1‡Xn iˆ1x2 i/C32! a2ˆXn iˆ1xiyi: …13:37† These equations can be solved for a1anda2. Problems 13.1. Given six points …ÿ1;0†,…ÿ0:8;2†,…ÿ0:6;1†,…ÿ0:4;ÿ1†,…ÿ0:2;0†;and …0;ÿ4†, determine a smooth function yˆf…x†such that yiˆf…xi†: 13.2. Find an approximate value of the real root of xÿtanxˆ0 near xˆ3=2: 13.3. Find the angle subtended at the center of a circle by an arc whose length is double the length of the chord. 13.4. Use Newton’s method to solve ex2ÿx3‡3xÿ4ˆ0; with x0ˆ0 and /C104ˆ0:001: 478NUMERICAL METHODS 13.5. Use Newton’s method to find a solution of sin…x3‡2†ˆ1=x; with x0ˆ1a n d /C104ˆ0:001. 13.6. Approximate the following integrals using the rectangular rule, the trapezoidal rule, and Simpson’s rule, with nˆ2;4;10;20;50: (a)Z=2 0eÿx2sin…x2‡1†dx ; (b)Z 2p 0sin…x2†‡3xÿ2 x‡4dx ; (c)Z1 0dx 2ÿsin2x/C112 : 13.7 Show that the area under a parabola, as shown in Fig. 13.7, is given by Aˆ/C104 3…y1‡4y2‡y3†: 13.8. Using the improved Euler’s method, find the value of ywhen xˆ0:2o n the integral curve of the equation y0ˆx2ÿ2ythrough the point xˆ0, yˆ1. 13.9. Using Taylor’s method, find correct to four places of decimals values of y corresponding to xˆ0:2 and xˆÿ0:2 for the solution of the di/C128erential equation dy=dxˆxÿy2=10; with the initial condition yˆ1 when xˆ0. 13.10. Using the Runge–Kutta method and /C104ˆ0:1, solve y0ˆx2ÿsin…y2†;x0ˆ1 and y0ˆ4:7: 13.11. Using the Runge–Kutta method and /C104ˆ0:1, solve y0ˆyeÿx2;x0ˆ1 and y0ˆ3: 13.12. Using Taylor’s method, obtain the solution of the system y0ˆx‡u;u0ˆ1‡y withy…0†ˆÿ 1;u…0†ˆ1: . 13.13. Find to four places of decimals the solution between xˆ0a n d xˆ0:5o f the equations y0ˆ1 2…y‡u†;u0ˆ12…y2ÿu2†; with yˆuˆ1 when xˆ0. 479PROBLEMS 13.14. Find to three places of decimals a solution of the equation y00‡2xy0ÿ4yˆ0; with yˆy0ˆ1 when xˆ0: 13.15. Use Eqs. (13.36) and (13.37) to calculate the coecients in yˆa1‡a2xto fit the following data: …x;y†ˆ… 1;1:7†;…2;1:8†;…3;2:3†;…4;3:2†: 480NUMERICAL METHODS 14 Introduction to probability theory The theory of probability is so useful that it is required in almost every branch of science. In physics, it is of basic importance in quantum mechanics, kinetic theory, and thermal and statistical physics to name just a few topics. In this chapter the reader is introduced to some of the fundamental ideas that make probability theory so useful. We begin with a review of the definitions of probability, a brief discussion of the fundamental laws of probability, and methods of counting (some facts about permutations and combinations), probability distributions are then treated. A notion that will be used very often in our discussion is ‘equally likely’. This cannot be defined in terms of anything simpler, but can be explained and illu-strated with simple examples. For example, heads and tails are equally likely results in a spin of a fair coin; the ace of spades and the ace of hearts are equally likely to be drawn from a shu/C130ed deck of 52 cards. Many more examples can be given to illustrate the concept of ‘equally likely’. /C65 definition of probabilit/C121 Now a question that arises naturally is that of how shall we measure the probability that a particular case (or outcome) in an experiment (such as the throw of dice or the draw of cards) out of many equally likely cases that will occur. Let us flip a coin twice, and ask the question: what is the probability of itcoming down heads at least once. There are four equally likely results in flipping a coin twice: /C72/C72,/C72T,T/C72/C44 TT , where /C72stands for head and Tfor tail. Three of the four results are favorable to at least one head showing, so the probability ofgetting one head is 3/4. In the example of drawn cards, what is the probability of drawing the ace of spades/C63 Obviously there is one chance out of 52, and the probability, accordingly, is 1/52. On the other hand, the probability of drawing an 481 ace is four times as great ÿ4=52, for there are four aces, equally likely. Reasoning in this way, we are led to give the notion of probability the following definition: If there are /C78mutually exclusive, collective exhaustive, and equally likely outcomes of an experiment, and nof these are favorable to an event A, then the probability /C112…A†of an event Aisn=/C78:/C112ˆn=/C78,o r /C112…A†ˆnumber of outcomes favorable to A total number of results: …14:1† We have made no attempt to predict the result, just to measure it. The definition of probability given here is often called a posteriori probability. The terms exclusive and exhaustive need some attention. Two events are said to be mutually exclusive if they cannot both occur together in a single trial; and the term collective exhaustive means that all possible outcomes or results are enum- erated in the /C78outcomes. If an event is certain not to occur its probability is zero, and if an event is certain to occur, then its probability is 1. Now if pis the probability that an event will occur, then the probability that it will fail to occur is 1 ÿ/C112, and we denote it by/C113: /C113ˆ1ÿ/C112: …14:2† Ifpis the probability that an event will occur in an experiment, and if the experiment is repeated Mtimes, then the expected number of times the event will occur is Mp. For suciently large M,Mpis expected to be close to the actual number of times the event will occur. For example, the probability of a headappearing when tossing a coin is 1/2, the expected number of times heads appear is 41=2 or 2. Actually, heads will not always appear twice when a coin is tossed four times. But if it is tossed 50 times, the number of heads that appear will, on theaverage, be close to 25 …501=2ˆ25). Note that closeness is computed on a percentage basis: 20 is 20/C37 of 25 away from 25 while 1 is 50/C37 of 2 away from 2. /C83ample space The equally likely cases associated with an experiment represent the possible out-comes. For example, the 36 equally likely cases associated with the throw of a pairof dice are the 36 ways the dice may fall, and if 3 coins are tossed, there are 8 equally likely cases corresponding to the 8 possible outcomes. A list or set that consists of all possible outcomes of an experiment is called a sample space and each individual outcome is called a sample point (a point of the sample space). The outcomes composing the sample space are required to be mutually exclusive. As an example, when tossing a die the outcomes ‘an even number shows’ and 482INTRODUCTION TO PROBABILITY THEORY ‘number 4 shows’ cannot be in the same sample space. Often there will be more than one sample space that can describe the outcome of an experiment but there isusually only one that will provide the most information. In a throw of a fair die, one sample space is the set of all possible outcomes /C1231, 2, 3, 4, 5, 6/C125, and another could be /C123even/C125 or /C123odd/C125. A finite sample space is one that has only a finite number of points. The points of the sample space are weighted according to their probabilities. To see this, let the points have the probabilities /C112 1;/C1122;...;/C112/C78 with /C1121‡/C1122‡‡ /C112/C78ˆ1: Suppose the first nsample points are favorable to another event A. Then the probability of Ais defined to be /C112…A†ˆ/C1121‡/C1122‡‡ /C112n: Thus the points of the sample space are weighted according to their probabilities. If each point has the sample probability 1/ n, then /C112…A†becomes /C112…A†ˆ1 /C78‡1 /C78‡‡1 /C78ˆn /C78 and this definition is consistent with that given by Eq. (14.1). A sample space with constant probability is called uniform. Non-uniform sam- ple spaces are more common. As an example, let us toss four coins and count the number of heads. An appropriate sample space is composed of the outcomes 0 heads ;1 head ;2 heads ;3 heads ;4 heads ; with respective probabilities, or weights 1=16;4=16;6=16;4=16;1=16: The four coins can fall in 2 222ˆ24, or 16 ways. They give no heads (all land tails) in only one outcome, and hence the required probability is 1/16. There are four ways to obtain 1 head: a head on the first coin or on the second coin, and so on. This gives 4/16. Similarly we can obtain the probabilities for the other cases. We can also use this simple example to illustrate the use of sample space. What is the probability of getting at least two heads/C63 Note that the last three samplepoints are favorable to this event, hence the required probability is given by 6 16‡4 16‡1 16ˆ11 16: 483SAMPLE SPACE Methods of counting In many applications the total number of elements in a sample space or in an event needs to be counted. A fundamental principle of counting is this: if one thing can be done in ndi/C128erent ways and another thing can be done in mdi/C128erent ways, then both things can be done together or in succession in mndi/C128erent ways. As an example, in the example of throwing a pair of dice cited above, there are 36 equally like outcomes: the first die can fall in six ways, and for each of these the second die can also fall in six ways. The total number of ways is 6‡6‡6‡6‡6‡6ˆ66ˆ36 and these are equally likely. Enumeration of outcomes can become a lengthy process, or it can become a practical impossibility. For example, the throw of four dice generates a samplespace with 6 4ˆ1296 elements. Some systematic methods for the counting are desirable. Permutation and combination formulas are often very useful. Permutations A permutation is a particular ordered selection. Suppose there are nobjects and r of these objects are arranged into rnumbered spaces. Since there are nways of choosing the first object, and after this is done there are nÿ1 ways of choosing the second object, ...;and finally nÿ…rÿ1†ways of choosing the rth object, it follows by the fundamental principle of counting that the number of di/C128erentarrangements or permutations is given by nPrˆn…nÿ1†…nÿ2†… nÿr‡1†: …14:3† where the product on the right-hand side has rfactors. We call nPrthe number of permutations of nobjects taken rat a time. When rˆn, we have nPnˆn…nÿ1†…nÿ2†1ˆn/C33: We can rewrite nPrin terms of factorials: nPrˆn…nÿ1†…nÿ2†… nÿr‡1† ˆn…nÿ1†…nÿ2†… nÿr‡1†…nÿr†21 …nÿr†21 ˆn/C33 …nÿr†/C33: When rˆn,w eh a v e nPnˆn/C33=…nÿn†/C33ˆn/C33=0/C33. This reduces to n/C33 if we have 0/C33ˆ1 and mathematicians actually take this as the definition of 0/C33. Suppose the nobjects are not all di/C128erent. Instead, there are n1objects of one kind (that is, indistinguishable from each other), n2that is of a second kind ;...;nk 484INTRODUCTION TO PROBABILITY THEORY of a kth kind so that n1‡n2‡‡ nkˆn. A natural question is that of how many distinguishable arrangements are there of these nobjects. Assuming that there are /C78di/C128erent arrangements, and each distinguishable arrangement appears n1/C33,n2/C33;...times, where n1/C33is the number of ways of arranging the n1objects, similarly for n2/C33;...;nk/C33:Then multiplying /C78byn1/C33n2/C33;...;nk/C33we obtain the number of ways of arranging the nobjects if they were all distinguishable, that is,nPnˆn/C33: /C78n1/C33n2/C33nk/C33ˆn/C33 or /C78ˆn/C33=…n1/C33n2/C33...nk/C33†: /C78is often written as nPn1n2:::nk, and then we have nPn1n2:::nkˆn/C33 n1/C33n2/C33nk/C33: …14:4† For example, given six coins: one penny, two nickels and three dimes, the number of permutations of these six coins is 6P123ˆ6/C33=1/C332/C333/C33ˆ60: /C67ombinations A permutation is a particular ordered selection. Thus 123 is a di/C128erent permuta-tion from 231. In many problems we are interested only in selecting objects with- out regard to order. Such selections are called combinations. Thus 123 and 231 are now the same combination. The notation for a combination is nCrwhich means the number of ways in which robjects can be selected from nobjects without regard to order (also called the combination of nobjects taken rat a time). Among the nPrpermutations there are r/C33 that give the same combination. Thus, the total number of permutations of ndi/C128erent objects selected rat a time is r/C33nCrˆnPrˆn/C33 …nÿr†/C33: Hence, it follows that nCrˆn/C33 r/C33…nÿr†/C33: …14:5† It is straightforward to show that nCrˆn/C33 r/C33…nÿr†/C33ˆn/C33 ‰nÿ…nÿr†Š/C33…nÿr†/C33ˆnCnÿr: nCris often written as nCrˆn r : 485METHODS OF COUNTING The numbers (14.5) are often called binomial coecients because they arise in the binomial expansion …x‡y†nˆxn‡n 1 xnÿ1y‡n2 x nÿ2y2‡‡nn y n: When nis very large a direct evaluation of n/C33 is impractical. In such cases we use Stirling’s approximate formula n/C33/C25 2np nneÿn: The ratio of the left hand side to the right hand side approaches 1 as n!1 . For this reason the right hand side is often called an asymptotic expansion of the left hand side. Fundamental probabilit/C121 theorems So far we have calculated probabilities by directly making use of the definitions; itis doable but it is not always easy. Some important properties of probabilities will help us to cut short our computation works. These important properties are often described in the form of theorems. To present these important theorems, let us consider an experiment, involving two events Aand B, with /C78equally likely outcomes and let n 1ˆnumber of outcomes in which Aoccurs ;but not B; n2ˆnumber of outcomes in which Boccurs ;but not A; n3ˆnumber of outcomes in which both AandBoccur ; n4ˆnumber of outcomes in which neither AnorBoccurs : This covers all possibilities, hence n1‡n2‡n3‡n4ˆ/C78: The probabilities of AandBoccurring are respectively given by P…A†ˆn1‡n3 /C78; P…B†ˆn2‡n3 /C78; …14:6† the probability of either AorB(or both) occurring is P…A‡B†ˆn1‡n2‡n3 /C78; …14:7† and the probability of both AandBoccurring successively is P…AB†ˆn3 /C78: …14:8† Let us rewrite P…AB†as P…AB†ˆn3 /C78ˆn1‡n3 /C78n3 n1‡n3: 486INTRODUCTION TO PROBABILITY THEORY Now …n1‡n3†=/C78isP…A†by definition. After Ahas occurred, the only possible cases are the …n1‡n3†cases favorable to A. Of these, there are n3cases favorable toB, the quotient n3=…n1‡n3†represents the probability of Bwhen it is known that Aoccurred, PA…B†. Thus we have P…AB†ˆP…A†PA…B†: …14:9† This is often known as the theorem of joint (or compound) probability. In words, the joint probability (or the compound probability) of AandBis the product of the probability that Awill occur times the probability that Bwill occur if Adoes. PA…B†is called the conditional probability of Bgiven A(that is, given that Ahas occurred). To illustrate the theorem of joint probability (14.9), we consider the probability of drawing two kings in succession from a shu/C130ed deck of 52 playing cards. The probability of drawing a king on the first draw is 4/52. After the first king has been drawn, the probability of drawing another king from the remaining 51 cards is 3/51, so that the probability of two kings is 4 523 51ˆ1 221: If the events Aand Bare independent, that is, the information that Ahas occurred does not influence the probability of B, then PA…B†ˆP…B†and the joint probability takes the form P…AB†ˆP…A†P…B†;for independent events : …14:10† As a simple example, let us toss a coin and a die, and let Abe the event ‘head shows’ and Bis the event ‘4 shows.’ These events are independent, and hence the probability that 4 and a head both show is P…AB†ˆP…A†P…B†ˆ… 1=2†…1=6†ˆ1=12: Theorem (14.10) can be easily extended to any number of independent events A;B;C;...: Besides the theorem of joint probability, there is a second fundamental relation- ship, known as the theorem of total probability. To present this theorem, let us go back to Eq. (14.4) and rewrite it in a slightly di/C128erent form P…A‡B†ˆn1‡n2‡n3 /C78 ˆn1‡n2‡2n3ÿn3 /C78ˆ…n1‡n3†‡…n2‡n3†ÿn3 /C78 ˆn1‡n3 /C78‡n2‡n3 /C78ÿn3 /C78ˆP…A†‡P…B†ÿP…AB†; P…A‡B†ˆP…A†‡P…B†ÿP…AB†: …14:11† 487FUNDAMENTAL PROBABILITY THEOREMS This theorem can be represented diagrammatically by the intersecting points sets Aand Bshown in Fig. 14.1. To illustrate this theorem, consider the simple example of tossing two dice and find the probability that at least one die gives2. The probability that both give 2 is 1/36. The probability that the first die gives 2is 1/6, and similarly for the second die. So the probability that at least one gives 2 is P…A‡B†ˆ1=6‡1=6ÿ1=36ˆ11=36: For mutually exclusive events, that is, for events A,Bwhich cannot both occur, P…AB†ˆ0 and the theorem of total probability becomes P…A‡B†ˆP…A†‡P…B†; for mutually exclusive events : …4:12† For example, in the toss of a die, ‘4 shows’ (event A) and ‘5 shows’ (event B)a r e mutually exclusive, the probability of getting either 4 or 5 is P…A‡B†ˆP…A†‡P…B†ˆ1=6‡1=6ˆ1=3: The theorems of total and joint probability for uniform sample spaces estab- lished above are also valid for arbitrary sample spaces. Let us consider a finite sample space, its events /C69 iare so numbered that /C691;/C692;...;/C69jare favorable to A;/C69j‡1;...;/C69kare favorable to both AandB, and /C69k‡1;...;/C69mare favorable to B only. If the associated probabilities are /C112i, then Eq. (14.11) is equivalent to the identity /C1121‡‡ /C112mˆ…/C1121‡‡ /C112j‡/C112j‡1‡‡ /C112k† ‡…/C112j‡1‡‡ /C112k‡/C112k‡1‡‡ /C112m†ÿ…/C112j‡1‡‡ /C112m†: The sums within the three parentheses on the right hand side represent, respec-tively, P…A†P…B†;andP…AB†by definition. Similarly, we have P…AB†ˆ/C112 j‡1‡‡ /C112k ˆ…/C1121‡‡ /C112k†/C112j‡1 /C1121‡‡ /C112k‡‡/C112k /C1121‡‡ /C112k ˆP…A†PA…B†; which is Eq. (14.9). 488INTRODUCTION TO PROBABILITY THEORY Figure 14.1. /C82andom /C118ariables and probabilit/C121 distributions As demonstrated above, simple probabilities can be computed from elementary considerations. We need more ecient ways to deal with probabilities of whole classes of events. For this purpose we now introduce the concepts of random variables and a probability distribution. Random variables A process such as spinning a coin or tossing a die is called random since it isimpossible to predict the final outcome from the initial state. The outcomes of a random process are certain numerically valued variables that are often called random variables. For example, suppose that three dimes are tossed at the same time and we ask how many heads appear. The answer will be 0, 1, 2, or 3 heads, and the sample space Shas 8 elements: SˆfTTT ;HTT ;THT ;TTH ;HHT ;HTH ;THH ;HHH g: The random variable Xin this case is the number of heads obtained and it assumes the values 0;1;1;1;2;2;2;3: For instance, Xˆ1 corresponds to each of the three outcomes: HTT ;THT ;TTH . That is, the random variable Xcan be thought of as a function of the number of heads appear. A random variable that takes on a finite or countable infinite number of values (that is it has as many values as the natural numbers 1 ;2;3;...) is called a discrete random variable while one that takes on a non-countable infinite number ofvalues is called a non-discrete or continuous random variable. Probability distributions A random variable, as illustrated by the simple example of tossing three dimes atthe same time, is a numerical-valued function defined on a sample space. In symbols, X…s i†ˆxi iˆ1;2;...;n; …14:13† where siare the elements of the sample space and xiare the values of the random variable X. The set of numbers xican be finite or infinite. In terms of a random variable we will write P…Xˆxi†as the probability that the random variable Xtakes the value xi, and P…X<xi†as the probability that the random variable takes values less than xi, and so on. For simplicity, we often write P…Xˆxi†as/C112i. The pairs …xi;/C112i†foriˆ1;2;3;...define the probability 489RANDOM VARIABLES AND PROBABILITY DISTRIBUTIONS distribution or probability function for the random variable X. Evidently any probability distribution /C112ifor a discrete random variable must satisfy the follow- ing conditions: (i)0/C112i1; (ii) the sum of all the probabilities must be unity (certainty),P i/C112iˆ1: Expectation and variance The expectation or expected value or mean of a random variable is defined in terms of a weighted average of outcomes, where the weighting is equal to the probability /C112iwith which xioccurs. That is, if Xis a random variable that can take the values x1;x2;...;with probabilities /C1121;/C1122;...;then the expectation or expected value /C69…X†is defined by /C69…X†ˆ/C1121x1‡/C1122x2‡ˆX i/C112ixi: …14:14† Some authors prefer to use the symbol /C22for the expectation value /C69…X†. For the three dimes tossed at the same time, we have xiˆ01 2 3 /C112iˆ1=83=83=81=8 and /C69…X†ˆ1 80‡381‡382‡183ˆ32: We often want to know how much the individual outcomes are scattered away from the mean. A quantity measure of the spread is the di/C128erence Xÿ/C69…X†and this is called the deviation or residual. But the expectation value of the deviations is always zero: /C69…Xÿ/C69…X†† ˆX i…xiÿ/C69…X††/C112iˆX ixi/C112iÿ/C69…X†X i/C112i ˆ/C69…X†ÿ/C69…X†1ˆ0: This should not be particularly surprising; some of the deviations are positive, and some are negative, and so the mean of the deviations is zero. This means that the mean of the deviations is not very useful as a measure of spread. We get around the problem of handling the negative deviations by squaring each deviation, thereby obtaining a quantity that is always positive. Its expectation value is calledthe variance of the set of observations and is denoted by  2 2ˆ/C69‰…Xÿ/C69…X††2Šˆ/C69‰…Xÿ/C22†2Š: …14:15† 490INTRODUCTION TO PROBABILITY THEORY The square root of the variance, , is known as the standard deviation, and it is always positive. We now state some basic rules for expected values. The proofs can be found in any standard textbook on probability and statistics. In the following cis a con- stant, Xand/C89are random variables, and /C104…X†is a function of X: (1)/C69…cX†ˆc/C69…X†; (2)/C69…X‡Y†ˆ/C69…X†‡/C69…Y†; (3)/C69…XY†ˆ/C69…X†/C69…Y†(provided Xand/C89are independent); (4)/C69…/C104…X†† ˆP i/C104…xi†/C112i(for a finite distribution). /C83pecial probabilit/C121 distributions We now consider some special probability distributions in which we will use all the things we have learned so far about probability. /C84he binomial distribution Before we discuss the binomial distribution, let us introduce a term, the Bernoullitrials. Consider an experiment such as spinning a coin or throw a die repeatedly. Each spin or toss is called a trial. In any single trial there will be a probability p associated with a particular event (or outcome). If pis constant throughout (that is, does not change from one trial to the next), such trials are then said to be independent and are known as Bernoulli trials. Now suppose that we have nindependent events of some kind (such as tossing a coin or die), each of which has a probability pof success and probability of /C113ˆ…1ÿ/C112†of failure. What is the probability that exactly mof the events will succeed/C63 If we select mevents from n, the probability that these mwill succeed and all the rest …nÿm†will fail is /C112 m/C113nÿm. We have considered only one particular group or combination of mevents. How many combinations of mevents can be chosen from n/C63 It is the number of combinations of nthings taken mat a time: nCm. Thus the probability that exactly mevents will succeed from a group of nis f…m†ˆP…Xˆm†ˆ nCm/C112m/C113…nÿm† ˆn/C33 m/C33…nÿm†/C33/C112m/C113…nÿm†: …14:16† This discrete probability function (14.16) is called the binomial distribution for X, the random variable of the number of successes in the ntrials. It gives the prob- ability of exactly msuccesses in nindependent trials with constant probability p. Since many statistical studies involve repeated trials, the binomial distribution has great practical importance. 491SPECIAL PROBABILITY DISTRIBUTIONS Why is the discrete probability function (14.16) called the binomial distribu- tion/C63 Since for mˆ0;1;2;...;nit corresponds to successive terms in the binomial expansion …/C113‡/C112†nˆ/C113n‡nC1/C113nÿ1/C112‡nC2/C113nÿ2/C1122‡‡ /C112nˆXn mˆ0nCm/C112m/C113nÿm: To illustrate the use of the binomial distribution (14.16), let us find the prob- ability that a one will appear exactly 4 times if a die is thrown 10 times. Here nˆ10,mˆ4,/C112ˆ1=6, and /C113ˆ…1ÿ/C112†ˆ5=6. Hence the probability is f…4†ˆP…Xˆ4†ˆ10/C33 4/C336/C331 6456 6 ˆ0:0543: A few examples of binomial distributions, computed from Eq. (14.16), are shown in Figs. 14.2, and 14.3 by means of histograms. One of the key requirements for a probability distribution is that Xn mˆ/C111f…m†ˆXn mˆ/C111nCm/C112m/C113nÿmˆ1: …14:17† To show that this is in fact the case, we note that Xn mˆ/C111nCm/C112m/C113nÿm 492INTRODUCTION TO PROBABILITY THEORY Figure 14.2. The distribution is symmetric about mˆ10: is exactly equal to the binomial expansion of …/C113‡/C112†n. But here /C113‡/C112ˆ1, so …/C113‡/C112†nˆ1 and our proof is established. The mean (or average) number of successes, m, is given by mˆXn mˆ0mnCm/C112m…1ÿ/C112†nÿm: …14:18† The sum ranges from mˆ0t onbecause in every one of the sets of trials the same number of successes between 0 and nmust occur. It is similar to Eq. (14.17); the di/C128erence is that the sum in Eq. (14.18) contains an extra factor n. But we can convert it into the form of the sum in Eq. (14.17). Di/C128erentiating both sides of Eq. (14.17) with respect to p, which is legitimate as the equation is true for all p between 0 and 1, gives X nCm‰m/C112mÿ1…1ÿ/C112†nÿmÿ…nÿm†/C112m…1ÿ/C112†nÿmÿ1Šˆ0; where we have dropped the limits on the sum, remembering that mranges from 0 ton. The last equation can be rewritten as X mnCm/C112mÿ1…1ÿ/C112†nÿmˆX …nÿm†nCm/C112m…1ÿ/C112†nÿmÿ1 ˆnX nCm/C112m…1ÿ/C112†nÿmÿ1ÿX mnCm/C112m…1ÿ/C112†nÿmÿ1 or X mnCm‰/C112mÿ1…1ÿ/C112†nÿm‡/C112m…1ÿ/C112†nÿmÿ1ŠˆnX nCm/C112m…1ÿ/C112†nÿmÿ1: 493SPECIAL PROBABILITY DISTRIBUTIONS Figure 14.3. The distribution favors smaller value of m. Now multiplying both sides by /C112…1ÿ/C112†we get X mnCm‰…1ÿ/C112†/C112m…1ÿ/C112†nÿm‡/C112m‡1…1ÿ/C112†nÿmŠˆn/C112X nCm/C112m…1ÿ/C112†nÿm: Combining the two terms on the left hand side, and using Eq. (14.17) in the right hand side we have X mnCm/C112m…1ÿ/C112†nÿmˆX mf…m†ˆn/C112: …14:19† Note that the left hand side is just our original expression for m, Eq. (14.18). Thus we conclude that mˆn/C112 …14:20† for the binomial distribution. The variance 2is given by 2ˆX …mÿm†2f…m†ˆX …mÿn/C112†2f…m†; …14:21† here we again drop the summation limits for convenience. To evaluate this sumwe first rewrite Eq. (14.21) as  2ˆX …m2ÿ2mn/C112‡n2/C1122†f…m† ˆX m2f…m†ÿ2n/C112X mf…m†‡n2/C1122X f…m†: This reduces to, with the help of Eqs. (14.17) and (14.19), 2ˆX m2f…m†ÿ…n/C112†2: …14:22† To evaluate the first term on the right hand side, we first di/C128erentiate Eq. (14.19): X mnCm‰m/C112mÿ1…1ÿ/C112†nÿmÿ…nÿm†/C112m…1ÿ/C112†nÿmÿ1Šˆ/C112; then multiplying by /C112…1ÿ/C112†and rearranging terms as before X m2 nCm/C112m…1ÿ/C112†nÿmÿn/C112X mnCm/C112m…1ÿ/C112†nÿmˆn/C112…1ÿ/C112†: By using Eq. (14.19) we can simplify the second term on the left hand side andobtain X m 2 nCm/C112m…1ÿ/C112†nÿmˆ…n/C112†2‡n/C112…1ÿ/C112† or X m2f…m†ˆn/C112…1ÿ/C112‡n/C112†: Inserting this result back into Eq. (14.22), we obtain 2ˆn/C112…1ÿ/C112‡n/C112†ÿ…n/C112†2ˆn/C112…1ÿ/C112†ˆn/C112/C113; …14:23† 494INTRODUCTION TO PROBABILITY THEORY and the standard deviation ; ˆn/C112/C113p: …14:24† Two di/C128erent limits of the binomial distribution for large nare of practical importance: (1) n!1 and/C112!0 in such a way that the product n/C112ˆremains constant; (2) both nandpnare large. The first case will result a new distribution, the Poisson distribution, and the second cases gives us the Gaussian (or Laplace) distribution. /C84he Poisson distribution Now n/C112ˆ;so/C112ˆ=n. The binomial distribution (14.16) then becomes f…m†ˆP…Xˆm†ˆn/C33 m/C33…nÿm†/C33 nm 1ÿ nnÿm ˆn…nÿ1†…nÿ2†… nÿm‡1† m/C33nmm1ÿ nnÿm ˆ1ÿ1 n 1ÿ2n 1ÿmÿ1 nm m/C331ÿ nnÿm :…14:25† Now as n!1 , 1ÿ1n 1ÿ2n 1ÿmÿ1 n !1; while 1ÿ nnÿm ˆ1ÿ nn 1ÿ nÿm !eÿÿ 1…† ˆ eÿ; where we have made use of the result lim n!11‡ nn ˆe : It follows that Eq. (14.25) becomes f…m†ˆP…Xˆm†ˆmeÿ m/C33: …14:26† This is known as the Poisson distribution. Note thatP1 mˆ0P…Xˆm†ˆ1;as it should. 495SPECIAL PROBABILITY DISTRIBUTIONS The Poisson distribution has the mean /C69…X†ˆX1 mˆ0mmeÿ m/C33ˆX1 mˆ1meÿ …mÿ1†/C33ˆX1 mˆ0meÿ m/C33 ˆeÿX1 mˆ0m m/C33ˆeÿeˆ; …14:27† where we have made use of the result X1 mˆ0m m/C33ˆe: The variance 2of the Poisson distribution is 2ˆVar…X†ˆ/C69‰…Xÿ/C69…X††2Šˆ/C69…X2†ÿ‰/C69…X†Š2 ˆX1 mˆ0m2meÿ m/C33ÿ2ˆeÿX1 mˆ1mm …mÿ1†/C33ÿ2 ˆeÿd deÿ ÿ2ˆ: …14:28† To illustrate the use of the Poisson distribution, let us consider a simple exam- ple. Suppose the probability that an individual su/C128ers a bad reaction from a flu injection is 0.001; what is the probability that out of 2000 individuals ( a) exactly 3, (b) more than 2 individuals will su/C128er a bad reaction/C63 Now Xdenotes the number of individuals who su/C128er a bad reaction and it is binomially distributed. However, we can use the Poisson approximation, because the bad reactions are assumed to be rare events. Thus P…Xˆm†ˆmeÿ m/C33;with ˆm/C112ˆ…2000†…0:001†ˆ2/C58 (a)P…Xˆ3†ˆ23eÿ2 3/C33ˆ0:18; …b†P…X/C622†ˆ1ÿ‰P…Xˆ0†‡P…Xˆ1†‡P…Xˆ2†Š ˆ1ÿ20eÿ2 0/C33‡21eÿ2 1/C33‡22eÿ2 2/C33"# ˆ1ÿ5eÿ2ˆ0:323: An exact evaluation of the probabilities using the binomial distribution would require much more labor. 496INTRODUCTION TO PROBABILITY THEORY The Poisson distribution is very important in nuclear physics. Suppose that we have nradioactive nuclei and the probability for any one of these to decay in a given interval of time Tisp, then the probability that mnuclei will decay in the interval Tis given by the binomial distribution. However, nmay be a very large number (such as 1023), and pmay be the order of 10ÿ20, and it is impractical to evaluate the binomial distribution with numbers of these magnitudes. Fortunately, the Poisson distribution can come to our rescue. The Poisson distribution has its own significance beyond its connection with the binomial distribution and it can be derived mathematically from elementary con- siderations. In general, the Poisson distribution applies when a very large number of experiments is carried out, but the probability of success in each is very small,so that the expected number of successes is a finite number. /C84he Gaussian /C40or normal/C41 distribution The second limit of the binomial distribution that is of interest to us results when both nandpnare large. Clearly, we assume that m,n,a n d nÿmare large enough to permit the use of Stirling’s formula ( n/C33/C25 2np n neÿn). Replacing m/C33,n/C33, and (nÿm)/C33 by their approximations and after simplification, we obtain P…Xˆm†/C129n/C112 mmn/C113 nÿmnÿmn 2m…nÿm†/C114 : …14:29† The binomial distribution has the mean value np(see Eq. (14.20). Now let  denote the deviation of mfrom np; that is, ˆmÿn/C112. Then nÿmˆn/C113ÿ; and Eq. (14.29) becomes P…Xˆm†ˆ1 2n/C112/C1131‡=n/C112 …† 1ÿ=n/C112 …†/C112 1‡ n/C112ÿ…n/C112‡† 1ÿ n/C113ÿ…n/C113ˆ† or P…Xˆm†Aˆ1‡ n/C112ÿ…n/C112‡† 1ÿ n/C113ÿ…n/C113ÿ† ; where Aˆ 2n/C112/C113 1‡ n/C112 1ÿ n/C113/C115 : Then logP…Xˆm†A …† /C129 ÿ … n/C112‡†log 1 ‡=n/C112 …† ÿ … n/C113ÿ†log…1ÿ=n/C113†: 497SPECIAL PROBABILITY DISTRIBUTIONS Assuming jj<n/C112/C113, so that =n/C112jj <1 and =n/C113jj <1, this permits us to write the two convergent series log 1 ‡ n/C112 ˆ n/C112ÿ2 2n2/C1122‡3 3n3/C1123ÿ ; log 1 ÿ n/C113 ˆÿ n/C113ÿ2 2n2/C1132ÿ3 3n3/C1133ÿ : Hence logP…Xˆm†A …† /C129 ÿ2 2n/C112/C113ÿ3…/C1122ÿ/C1132† 23n2/C1122/C1132ÿ4…/C1123‡/C1133† 34n3/C1123/C1133ÿ : Now, if jjis so small in comparison with np/C113that we ignore all but the first term on the right hand side of this expansion and Acan be replaced by …2n/C112/C113†1=2, then we get the approximation formula P…Xˆm†ˆ12n/C112/C113p eÿ2=2n/C112/C113: …14:30† When ˆn/C112/C113p;Eq. (14.30) becomes f…m†ˆP…Xˆm†ˆ1 2p eÿ2=22: …14:31† This is called the Guassian, or normal, distribution. It is a very good approxima- tion even for quite small values of n. The Gaussian distribution is a symmetrical bell-shaped distribution about its mean /C22, and is a measure of the width of the distribution. Fig. 14.4 gives a comparison of the binomial distribution and the Gaussian approximation. The Gaussian distribution also has a significance far beyond its connection with the binomial distribution. It can be derived mathematically from elementary con- siderations, and is found to agree empirically with random errors that actually 498INTRODUCTION TO PROBABILITY THEORY Figure 14.4. occur in experiments. Everyone believes in the Gaussian distribution: mathe- maticians think that physicists have verified it experimentally and physiciststhink that mathematicians have proved it theoretically. One of the main uses of the Gaussian distribution is to compute the probability X m2 mˆm1f…m† that the number of successes is between the given limits m1andm2. Eq. (14.31) shows that the above sum may be approximated by a sum X 1  2p eÿ2=22…14:32† over appropriate values of . Since ˆmÿn/C112, the di/C128erence between successive values of is 1, and hence if we let zˆ=, the di/C128erence between successive values of ziszˆ1=. Thus Eq. (14.32) becomes the sum over z, X 1 2peÿz2=2z: …14:33† Asz!0, the expression (14.33) approaches an integral, which may be evalu- ated in terms of the function …z†ˆZz 01 2peÿz2=2dzˆ1 2pZ z 0eÿz2=2dz: …14:34† The function ……z†is related to the extensively tabulated error function, erf( z): erf…z†ˆ2pZz 0eÿz2dz;and …z†ˆ1 2erfz 2p : These considerations lead to the following important theorem, which we state without proof: If mis the number of successes in nindependent trials with con- stant probability p, the probability of the inequality z1mÿn/C112n/C112/C113p z2 …14:35† approaches the limit 1  2pZz2 z1eÿz2=2dzˆ…z2†ÿ…z1†… 14:36† asn!1 . This theorem is known as Laplace–de Moivre limit theorem. To illustrate the use of the result (14.36), let us consider the simple example of a die tossed 600 times, and ask what the probability is that the number of ones will 499SPECIAL PROBABILITY DISTRIBUTIONS be between 80 and 110. Now nˆ600, /C112ˆ1=6,/C113ˆ1ÿ/C112ˆ5=6, and mvaries from 80 to 110. Hence z1ˆ80ÿ100  100…5=6†/C112 ˆÿ2:19 and z1ˆ110ÿ100 100…5=6†/C112 ˆ1:09: The tabulated error function gives …z 2†ˆ …1:09†ˆ0:362; and …z1†ˆ …ÿ2:19†ˆÿ …2:19†ˆÿ 0:486; where we have made use of the fact that …ÿz†ˆÿ …z†;you can check this with Eq. (14.34). So the required probability is approximately given by 0:362ÿ… ÿ 0:486†ˆ0:848: /C67ontinuous distributions So far we have discussed several discrete probability distributions: since measure- ments are generally made only to a certain number of significant figures, the variables that arise as the result of an experiment are discrete. However, discretevariables can be approximated by continuous ones within the experimental error. Also, in some applications a discrete random variable is inappropriate. We now give a brief discussion of continuous variables that will be denoted by x. We shall see that continuous variables are easier to handle analytically. Suppose we want to choose a point randomly on the interval 0 x1, how shall we measure the probabilities associated with that event/C63 Let us dividethis interval …0;1†into a number of subintervals, each of length xˆ0:1 (Fig. 14.5), the point xis then equally likely to be in any of these subintervals. The probability that 0 :3<x<0:6, for example, is 0.3, as there are three favorable cases. The probability that 0 :32<x<0:64 is found to be 0 :64ÿ0:32ˆ0:32 when the interval is divided into 100 parts, and so on. From these we see thatthe probability for xto be in a given subinterval of (0, 1) is the length of that subinterval. Thus P…a<x<b†ˆbÿa; 0ab1: …14:37† 500INTRODUCTION TO PROBABILITY THEORY Figure 14.5. The variable xis said to be uniformly distributed on the interval 0 x1. Expression (14.37) can be rewritten as P…a<x<b†ˆZb adxˆZb a1dx: For a continuous variable it is customary to speak of the probability density, which in the above case is unity. More generally, a variable may be distributed with an arbitrary density f…x†. Then the expression f…z†dz measures approximately the probability that xis on the interval z<x<z‡dz: And the probability that xis on a given interval ( a;b)i s P…a<x<b†ˆZb af…x†dx …14:38† as shown in Fig. 14.6. The function f…x†is called the probability density function and has the proper- ties: (1)f…x†0; …ÿ1 <x<1†; (2)Z1 ÿ1f…x†dxˆ1;a real-valued random variable must lie between 1. The function /C70…x†ˆP…Xx†ˆZx ÿ1f…u†du …14:39† defines the probability that the continuous random variable Xis in the interval (ÿ1;x†, and is called the cumulative distributive function. If f…x†is continuous, then Eq. (14.39) gives /C700…x†ˆf…x† and we may speak of a probability di/C128erential d/C70…x†ˆf…x†dx: 501CONTINUOUS DISTRIBUTIONS Figure 14.6. By analogy with those for discrete random variables the expected value or mean and the variance of a continuous random variable Xwith probability density function f…x†are defined, respectively, to be: /C69…X†ˆ/C22ˆZ1 ÿ1xf…x†dx; …14:40† Var…X†ˆ2ˆ/C69……Xÿ/C22†2†ˆZ1 ÿ1…xÿ/C22†2f…x†dx: …14:41† /C84he Gaussian /C40or normal/C41 distribution One of the most important examples of a continuous probability distribution is the Gaussian (or normal) distribution. The density function for this distribution is given by f…x†ˆ1  2peÿ…xÿ/C22†2=22; ÿ1 <x<1; …14:42† where /C22andare the mean and standard deviation, respectively. The correspond- ing distribution function is /C70…x†ˆP…Xx†ˆ1 2pZ x ÿ1eÿ…uÿ/C22†2=22du: …14:43† The standard normal distribution has mean zero …/C22ˆ0†and standard devia- tion ( ˆ1) f…z†ˆ1 2peÿz2=2: …14:44† Any normal distribution can be ‘standardized’ by considering the substitution zˆ…xÿ/C22†=in Eqs. (14.42) and (14.43). A graph of the density function (14.44), known as the standard normal curve, is shown in Fig. 14.7. We have also indi- cated the areas within 1, 2 and 3 standard deviations of the mean (that is between zˆÿ1 and ‡1,ÿ2 and ‡2,ÿ3a n d ‡3): P…ÿ1Z1†ˆ1  2pZ1 ÿ1eÿz2=2dzˆ0:6827 ; P…ÿ2Z2†ˆ1 2pZ 2 ÿ2eÿz2=2dzˆ0:9545 ; P…ÿ3Z3†ˆ1 2pZ 3 ÿ3eÿz2=2dzˆ0:9973: 502INTRODUCTION TO PROBABILITY THEORY The above three definite integrals can be evaluated by making numerical approx- imations. A short table of the values of the integral /C70…x†ˆ1  2pZx 0eÿt2dtˆ1 21 2pZx ÿxeÿt2dt is included in Appendix 3. A more complete table can be found in Tables of /C78ormal Probability /C70unctions , National Bureau of Standards, Washington, DC, 1953. /C84he /C77ax/C119ell/C177Bolt/C122mann distribution Another continuous distribution that is very important in physics is the Maxwell– Boltzmann distribution f…x†ˆ4aa /C114 x2eÿax2;0x<1;a/C620; …14:45† where aˆm=2kT,mis the mass, Tis the temperature (K), kis the Boltzmann constant, and xis the speed of a gas molecule. Problems 14.1 If a pair of dice is rolled what is the probability that a total of 8 shows/C63 14.2 Four coins are tossed, and we are interested in the number of heads. What is the probability that there is an odd number of heads/C63 What is the prob-ability that the third coin will land heads/C63 503PROBLEMS Figure 14.7. 14.3 Two coins are tossed. A reliable witness tells us ‘at least 1 coin showed heads.’ What e/C128ect does this have on the uniform sample space/C63 14.4 The tossing of two coins can be described by the following sample space: Event no heads one head two head Probability 1/4 1/2 1/4 What happens to this sample space if we know at least one coin showed heads but have no other specific information/C63 14.5 Two dice are rolled. What are the elements of the sample space/C63 What is the probability that a total of 8 shows/C63 What is the probability that at least one 5 shows/C63 14.6 A vessel contains 30 black balls and 20 white balls. Find the probability of drawing a white ball and a black ball in succession from the vessel. 14.7 Find the number of di/C128erent arrangements or permutations consisting of three letters each which can be formed from the seven letters A/C44 B/C44 /C67/C44 /C68/C44 E/C44 /C70/C44 /C71. 14.8 It is required to sit five boys and four girls in a row so that the girls occupy the even seats. How many such arrangements are possible/C63 14.9 A balanced coin is tossed five times. What is the probability of obtaining three heads and two tails/C63 14.10 How many di/C128erent five-card hands can be dealt from a shu/C130ed deck of 52 cards/C63 What is the probability that a hand dealt at random consists of fivespades/C63 14.11 ( a) Find the constant term in the expansion of ( x 2‡1=x†12: (b) Evaluate 50/C33. 14.12 A box contains six apples of which two are spoiled. Apples are selected at random without replacement until a spoiled one is found. Find the probability distribution of the number of apples drawn from the box, and present this distribution graphically. 14.13 A fair coin is tossed six times. What is the probability of getting exactly two heads/C63 14.14 Suppose three dice are rolled simultaneously. What is the probability that two 5 sappear with the third face showing a di/C128erent number/C63 14.15 Verify thatP1 mˆ0P…Xˆm†ˆ1 for the Poisson distribution. 14.16 Certain processors are known to have a failure rate of 1.2/C37. There are shipped in batches of 150. What is the probability that a batch has exactly one defective processor/C63 What is the probability that it has two/C63 14.17 A Geiger counter is used to count the arrival of radioactive particles. Find: (a) the probability that in time tno particles will be counted; (b) the probability of exactly one count in time t. 504INTRODUCTION TO PROBABILITY THEORY 14.18 Given the density function f…x† f…x†ˆkx20<x<3 0 otherwise/C58( (a) find the constant k; (b) compute P…1<x<2†; (c) find the distribution function and use it to find P…1<x…2††: 505PROBLEMS Appendix 1 Preliminaries (review of fundamental concepts) This appendix is for those readers who need a review; a number of fundamental concepts or theorem will be reviewed without giving proofs or attempting to achieve completeness. We assume that the reader is already familiar with the classes of real numbers used in analysis. The set of positive integers (also known as natural numbers) 1, 2, ...;nadmits the operations of addition without restriction, that is, they can be added (and therefore multiplied) together to give other positiveintegers. The set of integers 0,1;2;...;nadmits the operations of addition and subtraction among themselves. /C82ational numbers are numbers of the form /C112=/C113, where pand/C113are integers and /C1136ˆ0. Examples of rational numbers are 2/3, ÿ10=7. This set admits the further property of division among its members. The set of irrational numbers includes all numbers which cannot be expressed as the quotient of two integers. Examples of irrational numbers are 2p ;11 3p ;and any number of the form a=bn/C112 , where aand bare integers which are perfect nth powers. The set of real numbers contains all the rationals and irrationals. The important property of the set of real numbers fxgis that it can be put into (1:1) cor- respondence with the set of points fPgof a line as indicated in Fig. A.1. The basic rules governing the combinations of real numbers are: commutative law: a‡bˆb‡a;abˆba; associative law: a‡…b‡c†ˆ…a‡b†‡c;a…bc†ˆ…ab†c; distributive law: a…b‡c†ˆab‡ac; index law amanˆam‡n;am=anˆamÿn…a6ˆ0†; where a;b;c, are algebraic symbols for the real numbers. Problem A1.1 Prove that 2p is an irrational number. 506 (Hint: Assume the contrary, that is, assume that 2p ˆ/C112=/C113, where pand /C113are positive integers having no common integer factor.) Inequalities Ifxandyare real numbers, x/C62ymeans that xis greater than y;a n d x<ymeans that xis less than y. Similarly, xyimplies that xis either greater than or equal toy. The following basic rules governing the operations with inequalities: (1) Multiplication by a constant: If x/C62y, then ax/C62ayifais a positive num- ber, and ax<ayifais a negative number. (2) Addition of inequalities: If x;y;u;/C118are real numbers, and if x/C62y, and u/C62/C118, than x‡u/C62y‡/C118. (3) Subtraction of inequalities: If x/C62y,a n d u/C62/C118, we cannot deduce that …xÿu†/C62…yÿ/C118†. Why/C63 It is evident that …xÿu†ÿ…yÿ/C118†ˆ …xÿy†ÿ…uÿ/C118†is not necessarily positive. (4) Multiplication of inequalities: If x/C62y, and u/C62/C118, and x;y;u;/C118areallposi- tive, then xu/C62y/C118. When some of the numbers are negative, then the result is not necessarily true. (5) Division of inequalities: x/C62yandu/C62/C118do not imply x=u/C62y=/C118. When we wish to consider the numerical value of the variable xwithout regard to its sign, we write jxjand read this as ‘absolute or mod x’. Thus the inequality jxjais equivalent to ax‡ a. Problem A1.2 Find the values of xwhich satisfy the following inequalities: (a)x3ÿ7x2‡21xÿ27/C620, (b)j7ÿ3xj<2, (c)5 5xÿ1/C622 2x‡1. (Warning: cross multiplying is not permitted.) Problem A1.3 Ifa1;a2;...;anand b1;b2;...;bnare any real numbers, prove Schwarz’s inequality: …a1b1‡a2b2‡‡ anbn†2…a2 1‡a22‡‡ ann†…b21‡b22‡‡ bnn†: 507INEQUALITIES Figure A1.1. Problem A1.4 Show that 1 2‡14‡18‡‡1 2nÿ11 for all positive integers n/C621: Ifx1;x2;...;xnarenpositive numbers, their arithmetic mean is defined by Aˆ1nX n kˆ1xkˆx1‡x2‡‡ xn n and their geometric mean by /C71ˆn/C89n kˆ1xk/C115 ˆx1x2xnnp; wherePand/C81are the summation and product signs. The harmonic mean /C72is sometimes useful and it is defined by 1 Hˆ1 nXn kˆ11 xkˆ1n1 x1‡1 x2‡‡1 xn : There is a basic inequality among the three means: A/C71H, the equality sign occurring when x1ˆx2ˆˆ xn. Problem A1.5 Ifx1andx2are two positive numbers, show that A/C71H. Functions We assume that the reader is familiar with the concept of functions and theprocess of graphing functions. A polynomial of degree nis a function of the form f…x†ˆ/C112 n…x†ˆa0xn‡a1xnÿ1‡a2xnÿ2‡‡ an …ajˆconstant ;a06ˆ0†: A polynomial can be di/C128erentiated and integrated. Although we have written ajˆconstant, they might still be functions of some other variable independent ofx. For example, tÿ3x3‡sintx2‡ tp x‡t is a polynomial function of x(of degree 3) and each of the as is a function of a certain variable t:a0ˆtÿ3;a1ˆsint;a2ˆt1=2;a3ˆt. The polynomial equation f…x†ˆ0 has exactly nroots provided we count repe- titions. For example, x3ÿ3x2‡3xÿ1ˆ0 can be written …xÿ1†3ˆ0 so that the three roots are 1, 1, 1. Note that here we have used the binomial theorem …a‡x†nˆan‡nanÿ1x‡n…nÿ1† 2/C33anÿ2x2‡‡ xn: 508APPENDI/C88 1 PRELIMINARIES A rational function is of the form f…x†ˆ/C112n…x†=/C113n…x†, where /C112n…x†and/C113n…x† are polynomials. A transcendental function is any function which is not algebraic, for example, the trigonometric functions sin x, cos x, etc., the exponential functions ex, the logarithmic functions log x, and the hyperbolic functions sinh x, cosh x, etc. The exponential functions obey the index law. The logarithmic functions are inverses of the exponential functions, that is, if axˆythen xˆlogay, where ais called the base of the logarithm. If aˆe, which is often called the natural base of logarithms, we denote logexby ln x, called the natural logarithm of x. The funda- mental rules obeyed by logarithms are ln…mn†ˆlnm‡lnn;ln…m=n†ˆlnmÿlnn;and ln m/C112ˆ/C112lnm: The hyperbolic functions are defined in terms of exponential functions as follows sinhxˆexÿeÿx 2; coshxˆex‡eÿx 2; tanhxˆsinhx coshxˆexÿeÿx ex‡ex; cothxˆ1 tanhxˆex‡eÿx exÿeÿx; sechxˆ1 coshxˆ2 ex‡eÿx; cosech xˆ1 sinhxˆ2 exÿeÿx: Rough graphs of these six functions are given in Fig. A1.2. Some fundamental relationships among these functions are as follows: cosh2xÿsinh2xˆ1;sech2x‡tanh2xˆ1;coth2xÿcosech2xˆ1; sinh…xy†ˆsinhxcoshycoshxsinhy; cosh…xy†ˆcoshxcoshysinhxsinhy; tanh…xy†ˆtanhxtanhy 1tanhxtanhy: 509FUNCTIONS Figure A1.2. Hyperbolic functions. Problem A1.6 Using the rules of exponents, prove that ln …mn†ˆlnm‡lnn: Problem A1.7 Prove that: …a†sin2xˆ1 2…1ÿcos 2x†;cos2xˆ12…1‡cos 2x†, and ( b)Acosx‡ Bsinxˆ A2‡B2p sin…x‡†, where tan ˆA=B Problem A1.8 Prove that: …a†cosh2xÿsinh2xˆ1, and ( b)2x‡tanh2xˆ1. Limits We are sometimes required to find the limit of a function f…x†asxapproaches some particular value : lim x! f…x†ˆ/C108: This means that if jxÿ jis small enough, jf…x†ÿ/C108jcan be made as small as we please. A more precise analytic description of lim x! f…x†ˆ/C108is the following: For any /C34/C620 (however small) we can always find a number /C17 (which, in general, depends upon /C34) such that f…x†ÿ/C108 jj </C34 whenever xÿ jj </C17. As an example, consider the limit of the simple function f…x†ˆ2ÿ1=…xÿ1†as x!2. Then lim x!2f…x†ˆ1 for if we are given a number, say /C34ˆ10ÿ3, we can always find a number /C17which is such that 2ÿ1 xÿ1 ÿ1<10ÿ3…A1:1† provided jxÿ2j</C17. In this case (A1.1) will be true if 1 =…xÿ1†/C621ÿ10ÿ3ˆ 0:999. This requires xÿ1<…0:999†ÿ1,o rxÿ2<…0:999†ÿ1ÿ1. Thus we need only take /C17ˆ…0:999†ÿ1ÿ1. The function f…x†is said to be continuous at if lim x! f…x†ˆ/C108.I ff…x†is con- tinuous at each point ofan interval such as axbora<xb, etc., it is said to be continuous in the interval (for example, a polynomial is continuous at all x). The definition implies that lim x! ÿ0f…x†ˆlimx! ‡0f…x†ˆf… †at all points of the interval ( a;b), but this is clearly inapplicable at the endpoints aandb.A t these points we define continuity by lim x!a‡0f…x†ˆf…a†and lim x!bÿ0f…x†ˆf…b†: 510APPENDI/C88 1 PRELIMINARIES A finite discontinuity may occur at xˆ . This will arise when limx! ÿ0f…x†ˆ/C1081, lim x! ÿ0f…x†ˆ/C1082, and /C10816ˆ/C1082. It is obvious that a continuous function will be bounded in any finite interval. This means that we can find numbers mandMindependent of xand such that mf…x†Mforaxb. Furthermore, we expect to find x0;x1such that f…x0†ˆmandf…x1†ˆM. The order of magnitude of a function is indicated in terms of its variable. Thus, ifxis very small, and if f…x†ˆa1x‡a2x2‡a3x3‡ (akconstant), its magni- tude is governed by the term in xand we write f…x†ˆO…x†. When a1ˆ0, we write f…x†ˆO…x2†, etc. When f…x†ˆO…xn†, then lim x!0ff…x†=xngis finite and/ or lim x!0ff…x†=xnÿ1gˆ0. A function f…x†is said to be di/C128erentiable or to possess a derivative at the point xif lim /C104!0‰f…x‡/C104†ÿf…x†Š=/C104exists. We write this limit in various forms df=dx;f0orDf, where Dˆd…†=dx. Most of the functions in physics can be successively di/C128erentiated a number of times. These successive derivatives are written as f0…x†;f00…x†;...;fn…x†;...;orDf;D2f;...;Dnf;...: Problem A1.9 Iff…x†ˆx2, prove that: ( a) lim x!2f…x†ˆ4, and …b†f…x†is continuous at xˆ2. Infinite series Infinite series involve the notion of sequence in a simple way. For example, 2p is irrational and can only be expressed as a non-recurring decimal 1 :414 ...:We can approximate to its value by a sequence of rationals, 1, 1.4, 1.41, 1.414, ...sayfang which is a countable set limit of anwhose values approach indefinitely close to2p . Because of this we say the limit of a nasntends to infinity exists and equals2p , and write lim n!1anˆ2p . In general, a sequence u 1;u2;...;fungis a function defined on the set of natural numbers. The sequence is said to have the limit lor to converge to l, if given any /C34/C620 there exists a number /C78/C620 such that junÿ/C108j</C34for all n/C62/C78, and in such case we write lim n!1unˆ/C108. Consider now the sums of the sequence fung snˆXn rˆ1urˆu1‡u2‡u3‡ ; …A:2† where ur/C620 for all r.I fn!1 , then (A.2) is an infinite series of positive terms. We see that the behavior of this series is determined by the behavior of the sequence fungas it converges or diverges. If lim n!1snˆs(finite) we say that (A.2) is convergent and has the sum s. When sn!1 asn!1 , we say that (A.2) is divergent. 511INFINITE SERIES Example A1.1. Show that the series X1 nˆ11 2nˆ1 2‡1 22‡1 23‡ is convergent and has sum sˆ1. Solution: Let snˆ12‡1 22‡1 23‡‡1 2n; then 12s nˆ1 22‡1 23‡‡1 2n‡1: Subtraction gives 1ÿ12 s nˆ12ÿ1 2n‡1ˆ121ÿ1 2n ; or snˆ1ÿ1 2n: Then since lim n!1snˆlimn!1…1ÿ1=2n†ˆ1, the series is convergent and has the sum sˆ1. Example A1.2. Show that the seriesP1 nˆ1…ÿ1†nÿ1ˆ1ÿ1‡1ÿ1‡ is divergent. Solution: Here snˆ0 or 1 according as nis even or odd. Hence lim n!1sndoes not exist and so the series is divergent. Example A1.3. Show that the geometric seriesP1 nˆ1arnÿ1ˆa‡ar‡ar2‡ ;where aandrare constants, ( a) converges to sˆa=…1ÿr†ifjrj<1;and ( b) diverges if jrj/C621. Solution: Let snˆa‡ar‡ar2‡‡ arnÿ1: Then rsnˆ ar‡ar2‡‡ arnÿ1‡arn: Subtraction gives …1ÿr†snˆaÿarnor snˆa…1ÿrn† 1ÿr: 512APPENDI/C88 1 PRELIMINARIES (a)I fjrj<1; lim n!1snˆlim n!1a…1ÿrn† 1ÿrˆa 1ÿr: (b)I fjrj/C621, lim n!1snˆlim n!1a…1ÿrn† 1ÿr does not exist. Example A1.4. Show that the pseriesP1 nˆ11=n/C112converges if /C112/C621 and diverges if /C1121. Solution: Using f…n†ˆ1=npwe have f…x†ˆ1=xpso that if p6ˆ1, Z1 1dx x/C112ˆlim M!1ZM 1xÿ/C112dxˆlim M!1x1ÿ/C112 1ÿ/C112/C12/C12/C12/C12M 1ˆlim M!1M1ÿ/C112 1ÿ/C112ÿ1 1ÿ/C112"# : Now if /C112/C621 this limit exists and the corresponding series converges. But if /C112<1 the limit does not exist and the series diverges. If/C112ˆ1 then Z1 1dx xˆlim M!1ZM 1dx xˆlim M!1lnx/C12/C12/C12/C12M 1ˆlim M!1lnM; which does not exist and so the corresponding series for /C112ˆ1 diverges. This shows that 1 ‡1 2‡13‡ diverges even though the nth term approaches zero. /C84ests for convergence There are several important tests for convergence of series of positive terms. Before using these simple tests, we can often weed out some very badly divergent series with the following preliminary test: If the terms of an infinite series do not tend to zero (that is, iflim n!1an6ˆ0†, the series diverges. If lim n!1anˆ0, we must test further : Four of the common tests are given below: /C67omparison test Ifun/C118n(alln), thenP1 nˆ1unconverges whenP1nˆ1/C118nconverges. If un/C118n(all n), thenP1 nˆ1undiverges whenP1nˆ1/C118ndiverges. 513INFINITE SERIES Since the behavior ofP1 nˆ1unis una/C128ected by removing a finite number of terms from the series, this test is true if un/C118norun/C118nfor all n/C62/C78. Note that n/C62/C78means from some term onward. Often, /C78ˆ1. Example A1.5 (a) Since 1 =…2n‡1†1=2nandP1=2nconverges,P1=…2n‡1†also converges. (b) Since 1 =lnn/C621=nandP1 nˆ21=ndiverges,P1nˆ21=lnnalso diverges. /C81uotient test Ifun‡1=un/C118n‡1=/C118n(alln), thenP1 nˆ1unconverges whenP1nˆ1/C118nconverges. And ifun‡1=un/C118n‡1=/C118n(alln), thenP1 nˆ1undiverges whenP1nˆ1/C118ndiverges. We can write unˆun unÿ1unÿ1 unÿ2u2 u1u1/C118n /C118nÿ1/C118nÿ1 /C118nÿ2/C1182 /C1181/C1181 so that un/C118nu1which proves the quotient test by using the comparison test. A similar argument shows that if un‡1=un/C118n‡1=/C118n(alln), thenP1nˆ1un diverges whenP1nˆ1/C118ndiverges. Example A1.6 Consider the series X1 nˆ14n2ÿn‡3 n3‡2n: For large n,…4n2ÿn‡3†=…n3‡2n†is approximately 4 =n. Taking unˆ…4n2ÿn‡3†=…n3‡2n†and /C118nˆ1=n, we have lim n!1un=/C118nˆ1. Now sinceP/C118nˆP1=ndiverges,Punalso diverges. /C68/C39Alembert/C39s ratio test:P1 nˆ1unconverges when un‡1=un<1 (all n/C78) and diverges when un‡1=un/C621. Write /C118nˆxnÿ1in the quotient test so thatP1 nˆ1/C118nis the geometric series with common ratio /C118n‡1=/C118nˆx. Then the quotient test proves thatP1 nˆ1unconverges when x<1 and diverges when x/C621: Sometimes the ratio test is stated in the following form: if lim n!1un‡1=unˆ/C26, thenP1nˆ1unconverges when /C26<1 and diverges when /C26/C621. Example A1.7 Consider the series 1‡1 2/C33‡1 3/C33‡‡1 n/C33‡ : 514APPENDI/C88 1 PRELIMINARIES Using the ratio test, we have un‡1 unˆ1 …n‡1†/C33/C41 n/C33ˆn/C33 …n‡1†/C33ˆ1 n‡1<1; so the series converges. Integral test. Iff…x†is positive, continuous and monotonic decreasing and is such that f…n†ˆunforn/C62/C78, thenPunconverges or diverges according as Z1 /C78f…x†dxˆlim M!1ZM /C78f…x†dx converges or diverges. We often have /C78ˆ1 in practice. To prove this test, we will use the following property of definite integrals: If in axb;f…x†/C103…x†, thenZb af…x†dxZb a/C103…x†dx. Now from the monotonicity of f…x†, we have un‡1ˆf…n‡1†f…x†f…n†ˆun; nˆ1;2;3;...: Integrating from xˆntoxˆn‡1 and using the above quoted property of definite integrals we obtain un‡1Zn‡1 nf…x†dxun; nˆ1;2;3;...: Summing from nˆ1t oMÿ1, u1‡u2‡‡ uMZM 1f…x†dxu1‡u2‡‡ uMÿ1: …A1:3† Iff…x†is strictly decreasing, the equality sign in (A1.3) can be omitted. If lim M!1RM 1f…x†dxexists and is equal to s, we see from the left hand inequal- ity in (A1.3) that u1‡u2‡‡ uMis monotonically increasing and bounded above by s, so thatPunconverges. If lim M!1RM 1f…x†dxis unbounded, we see from the right hand inequality in (A1.3) thatPundiverges. Geometrically, u1‡u2‡‡ uMis the total area of the rectangles shown shaded in Fig. A1.3, while u1‡u2‡‡ uMÿ1is the total area of the rectangles which are shaded and non-shaded. The area under the curve yˆf…x†from xˆ1 toxˆMis intermediate in value between the two areas given above, thus illus- trating the result (A1.3). 515INFINITE SERIES Example A1.8P1 nˆ11=n2converges since lim M!1RM 1dx=x2ˆlimM!1…1ÿ1=M†exists. Problem A1.10 Find the limit of the sequence 0.3, 0.33, 0 :333;...;and justify your conclusion. /C65lternating series test An alternating series is one whose successive terms are alternately positive and negative u1ÿu2‡u3ÿu4‡ :It converges if the following two conditions are satisfied: (a)jun‡1junjforn1; (b) lim n!1unˆ0 or lim n!1unjjˆ0 : The sum of the series to 2 Mis S2Mˆ…u1ÿu2†‡…u3ÿu4†‡‡… u2Mÿ1ÿu2M† ˆu1ÿ…u2ÿu3†ÿ…u4ÿu5†ÿÿ… u2Mÿ2ÿu2Mÿ1†ÿu2M: Since the quantities in parentheses are non-negative, we have S2M0; S2S4S6 S2Mu1: Therefore fS2Mgis a bounded monotonic increasing sequence and thus has the limit S. Also S2M‡1ˆS2M‡u2M‡1. Since lim M!1S2MˆSand lim M!1u2M‡1ˆ0 (for, by hypothesis, lim n!1unˆ0), it follows that lim M!1S2M‡1ˆ limM!1S2M‡limM!1u2M‡1ˆS‡0ˆS. Thus the partial sums of the series approach the limit Sand the series converges. 516APPENDI/C88 1 PRELIMINARIES Figure A1.3. Problem A1.11 Show that the error made in stopping after 2Mterms is less than or equal to u2M‡1. Example A1.9For the series 1ÿ1 2‡13ÿ14‡ˆX 1 nˆ1…ÿ1†nÿ1 n; we have unˆ… ÿ 1†n‡1=n;unjjˆ1=n;un‡1jj ˆ1=…n‡1†. Then for n1; un‡1jj unjj. Also we have lim n!1unjjˆ0. Hence the series converges. /C65bsolute and conditional convergence The seriesPunis called absolutely convergent ifPunjjconverges. IfPuncon- verges butPunjjdiverges, thenPunis said to be conditionally convergent. It is easy to show that ifPunjjconverges, thenPunconverges (in words, an absolutely convergent series is convergent). To this purpose, let SMˆu1‡u2‡‡ uMTMˆu1jj‡u2jj‡‡ uMjj; then SM‡TMˆ…u1‡u1jj †‡…u2‡u2jj †‡‡… uM‡uMjj † 2u1jj‡2u2jj‡‡ 2uMjj: SincePunjjconverges and since un‡unjj0, for nˆ1;2;3;...;it follows that SM‡TM is a bounded monotonic increasing sequence, and so limM!1…SM‡TM†exists. Also lim M!1TMexists (since the series is absolutely convergent by hypothesis), lim M!1SMˆlim M!1…SM‡TMÿTM†ˆ lim M!1…SM‡TM†ÿlim M!1TM must also exist and so the seriesPunconverges. The terms of an absolutely convergent series can be rearranged in any order, and all such rearranged series will converge to the same sum. We refer the reader to text-books on advanced calculus for proof. Problem A1.12 Prove that the series 1ÿ1 22‡1 32ÿ1 42‡1 52ÿ converges. 517INFINITE SERIES How do we test for absolute convergence/C63 The simplest test is the ratio test, which we now review, along with three others – Raabe’s test, the nth root test, and Gauss’ test. /C82atio test Let lim n!1jun‡1=unjˆL. Then the seriesPun: (a) converges (absolutely) if L<1; (b) diverges if L/C621; (c) the test fails if Lˆ1. Let us consider first the positive-termPun, that is, each term is positive. We must now prove that if lim n!1un‡1=unˆL<1, then necessarilyPunconverges. By hypothesis, we can choose an integer /C78so large that for all n/C78;…un‡1=un†<r, where L<r<1. Then u/C78‡1<ru/C78;u/C78‡2<ru/C78‡1<r2u/C78;u/C78‡3<ru/C78‡2<r3u/C78;etc: By addition u/C78‡1‡u/C78‡2‡ <u/C78…r‡r2‡r3‡ † and so the given series converges by the comparison test, since 0 <r<1. When the series has terms with mixed signs, we consider u1jj‡u2jj‡u3jj‡ , then by the above proof and because an absolutely convergent series isconvergent, it follows that if lim n!1un‡1=un jj ˆL<1, thenPunconverges absolutely. Similarly we can prove that if lim n!1un‡1=un jj ˆL/C621, the seriesPun diverges. Example A1.10 Consider the seriesP1 nˆ1…ÿ1†nÿ12n=n2. Here unˆ… ÿ 1†nÿ12n=n2. Then limn!1un‡1=un jj ˆlimn!12n2=…n‡1†2ˆ2. Since Lˆ2/C621, the series diverges. When the ratio test fails, the following three tests are often very helpful. /C82aabe/C39s test Let lim n!1n…1ÿun‡1=un jj † ˆ‘, then the seriesPun: (a) converges absolutely if ‘<1; (b) diverges if ‘/C621. The test fails if ‘ˆ1. The n throot test Let lim n!1 unjjn/C112 ˆR, then the seriesPun: 518APPENDI/C88 1 PRELIMINARIES (a) converges absolutely if R<1; (b) diverges if R/C621. The test fails if Rˆ1. /C71auss/C39 test If un‡1 un/C12/C12/C12/C12/C12/C12/C12/C12ˆ1ÿ/C71 n‡cn n2; where jcnj<Pfor all n/C62/C78, then the seriesPun: (a) converges (absolutely) if /C71/C621; (b) diverges or converges conditionally if /C711. Example A1.11 Consider the series 1 ‡2r‡r2‡2r3‡r4‡2r5‡ . The ratio test gives un‡1 un/C12/C12/C12/C12/C12/C12/C12/C12ˆ2rjj;nodd rjj=2;neven;/C26 which indicates that the ratio test is not applicable. We now try the nth root test:  u njjn/C112 ˆ 2rnjjn/C112 ˆ 2np rjj;nodd rnjjn/C112 ˆrjj; neven( and so lim n!1 unjjn/C112 ˆrjj. Thus if jrj<1 the series converges, and if jrj/C621 the series diverges. Example A1.12 Consider the series 1 32 ‡14 362 ‡147 3692 ‡‡147…3nÿ2† 369…3n† ‡ : The ratio test is not applicable, since lim n!1un‡1 un/C12/C12/C12/C12/C12/C12/C12/C12ˆlim n!1…3n‡1† …3n‡3†/C12/C12/C12/C12/C12/C12/C12/C122 ˆ1: But Raabe’s test gives lim n!1n1ÿun‡1 un/C12/C12/C12/C12/C12/C12/C12/C12 ˆlim n!1n1ÿ3n‡1 3n‡32() ˆ4 3/C621; and so the series converges. 519INFINITE SERIES Problem A1.13 Test for convergence the series 1 22 ‡13 242 ‡135 2462 ‡‡135…2nÿ1† 135…2n† ‡ : Hint: Neither the ratio test nor Raabe’s test is applicable (show this). Try Gauss’ test. /C83eries of functions and uniform con/C118ergence The series considered so far had the feature that undepended just on n. Thus the series, if convergent, is represented by just a number. We now consider series whose terms are functions of x;unˆun…x†. There are many such series of func- tions. The reader should be familiar with the power series in which the nth term is a constant times xn: S…x†ˆX1 nˆ0anxn: …A1:4† We can think of all previous cases as power series restricted to xˆ1. In later sections we shall see Fourier series whose terms involve sines and cosines, and other series in which the terms may be polynomials or other functions. In this section we consider power series in x. The convergence or divergence of a series of functions depends, in general, on the values of x. With xin place, the partial sum Eq. (A1.2) now becomes a function of the variable x: sn…x†ˆu1‡u2…x†‡‡ un…x†: …A1:5† as does the series sum. If we define S…x†as the limit of the partial sum S…x†ˆlim n!1sn…x†ˆX1 nˆ0un…x†; …A1:6† then the series is said to be convergent in the interval /C91 a,b/C93 (that is, axb), if for each /C34/C620 and each xin /C91a,b/C93 we can find /C78/C620 such that S…x†ÿsn…x† jj </C34 ; for all n/C78: …A1:7† If/C78depends only on /C34and not on x, the series is called uniformly convergent in the interval /C91 a,b/C93. This says that for our series to be uniformly convergent, it must be possible to find a finite /C78so that the remainder of the series after /C78terms,P1 iˆ/C78‡1ui…x†, will be less than an arbitrarily small /C34for all xin the given interval. The domain of convergence (absolute or uniform) of a series is the set of values ofxfor which the series of functions converges (absolutely or uniformly). 520APPENDI/C88 1 PRELIMINARIES We deal with power series in xexactly as before. For example, we can use the ratio test, which now depends on x, to investigate convergence or divergence of a series: r…x†ˆlim n!1un‡1 un/C12/C12/C12/C12/C12/C12/C12/C12ˆlim n!1/C12/C12/C12/C12an‡1xn‡1 anxn/C12/C12/C12/C12ˆxjjlim n!1an‡1 an/C12/C12/C12/C12/C12/C12/C12/C12ˆxjjr; rˆlim n!1an‡1 an/C12/C12/C12/C12/C12/C12/C12/C12; thus the series converges (absolutely) if jxjr<1o r jxj<Rˆ 1 rˆlim n!1an an‡1/C12/C12/C12/C12/C12/C12/C12/C12 and the domain of convergence is given by R/C58ÿR<x<R. Of course, we need to modify the above discussion somewhat if the power series does not contain every power of x. Example A1.13 For what value of xdoes the seriesP 1 nˆ1xnÿ1=n3nconverge/C63 Solution: Now unˆxnÿ1=n3n,a n d x6ˆ0 (if xˆ0 the series converges). We have lim n!1un‡1 un/C12/C12/C12/C12/C12/C12/C12/C12ˆlim n!1n 3…n‡1†xjjˆ1 3xjj: Then the series converges if jxj<3, and diverges if jxj/C623. Ifjxjˆ3, that is, xˆ3, the test fails. Ifxˆ3, the series becomesP1 nˆ11=3nwhich diverges. If xˆÿ3, the series becomesP1nˆ1…ÿ1†nÿ1=3nwhich converges. Then the interval of convergence is ÿ3x<3. The series diverges outside this interval. Furthermore, the series converges absolutely for ÿ3<x<3 and converges conditionally at xˆÿ3. As for uniform convergence, the most commonly encountered test is the Weierstrass Mtest: /C87eierstrass Mtest If a sequence of positive constants M1;M2;M3;...;can be found such that: ( a) Mnjun…x†jfor all xin some interval /C91 a,b/C93, and ( b)PMnconverges, thenPun…x†is uniformly and absolutely convergent in /C91 a,b/C93. The proof of this common test is direct and simple. SincePMnconverges, some number /C78exists such that for n/C78, X1 iˆ/C78‡1Mi</C34 : 521SERIES OF FUNCTIONS AND UNIFORM CONVERGENCE This follows from the definition of convergence. Then, with Mnjun…x†jfor all x in /C91a,b/C93, X1 iˆ/C78‡1ui…x†jj </C34 : Hence S…x†ÿsn…x† jj ˆ/C12/C12/C12/C12X1 iˆ/C78‡1ui…x†/C12/C12/C12/C12</C34 ; for all n/C78 and by definitionPu n…x†is uniformly convergent in /C91 a,b/C93. Furthermore, since we have specified absolute values in the statement of the Weierstrass Mtest, the seriesPun…x†is also seen to be absolutely convergent. It should be noted that the Weierstrass Mtest only provides a sucient con- dition for uniform convergence. A series may be uniformly convergent even when theMtest is not applicable. The Weierstrass Mtest might mislead the reader to believe that a uniformly convergent series must be also absolutely convergent, andconversely. In fact, the uniform convergence and absolute convergence are inde- pendent properties. Neither implies the other. A somewhat more delicate test for uniform convergence that is especially useful in analyzing power series is Abel’s test. We now state it without proof. /C65bel’s test If…a†u n…x†ˆanfn…x†, andPanˆA, convergent, and ( b) the functions fn…x†are monotonic ‰fn‡1…x†fn…x†Šand bounded, 0 fn…x†Mfor all xin /C91a,b/C93, thenPun…x†converges uniformly in /C91 a,b/C93. Example A1.14 Use the Weierstrass Mtest to investigate the uniform convergence of …a†X1 nˆ1cosnx n4; …b†X1 nˆ1xn n3=2; …c†X1 nˆ1sinnx n: Solution: (a) cos …nx†=n4/C12/C12/C12/C121=n4ˆMn. Then sincePMnconverges ( pseries with /C112ˆ4/C621), the series is uniformly and absolutely convergent for all xby theMtest. (b) By the ratio test, the series converges in the interval ÿ1x1 (or jxj1). For all xinjxj1;xn=n3=2/C12/C12/C12/C12/C12/C12ˆxjj n=n3=21=n3=2. Choosing Mnˆ1=n3=2, we see thatPMnconverges. So the given series converges uniformly for jxj1 by the Mtest. 522APPENDI/C88 1 PRELIMINARIES (c) sin …nx†=n=n jj  1=nˆMn. However,PMndoes not converge. The Mtest cannot be used in this case and we cannot conclude anything about the uniform convergence by this test. A uniformly convergent infinite series of functions has many of the properties possessed by the sum of finite series of functions. The following three are parti-cularly useful. We state them without proofs. (1) If the individual terms u n…x†are continuous in /C91 a,b/C93 and ifPun…x†con- verges uniformly to the sum S…x†in /C91a,b/C93, then S…x†is continuous in /C91 a,b/C93. Briefly, this states that a uniformly convergent series of continuous func-tions is a continuous function. (2) If the individual terms u n…x†are continuous in /C91 a,b/C93 and ifPun…x†con- verges uniformly to the sum S…x†in /C91a,b/C93, then Zb aS…x†dxˆX1 nˆ1Zb aun…x†dx or Zb aX1 nˆ1un…x†dxˆX1 nˆ1Zb aun…x†dx: Briefly, a uniform convergent series of continuous functions can be inte-grated term by term. (3) If the individual terms u n…x†are continuous and have continuous derivatives in /C91a,b/C93 and ifPun…x†converges uniformly to the sum S…x†whilePdun…x†=dxis uniformly convergent in /C91 a,b/C93, then the derivative of the series sum S…x†equals the sum of the individual term derivatives, d dxS…x†ˆX1 nˆ1d dxun…x†ord dxX1 nˆ1un…x†() ˆX1 nˆ1d dxun…x†: Term-by-term integration of a uniformly convergent series requires only con- tinuity of the individual terms. This condition is almost always met in physicalapplications. Term-by-term integration may also be valid in the absence of uni- form convergence. On the other hand term-by-term di/C128erentiation of a series is often not valid because more restrictive conditions must be satisfied. Problem A1.14 Show that the series sinx 13‡sin 2x 23‡‡sinnx n3‡ is uniformly convergent for ÿx. 523SERIES OF FUNCTIONS AND UNIFORM CONVERGENCE /C84heorems on po/C119er series When we are working with power series and the functions they represent, it is very useful to know the following theorems which we will state without proof. We will see that, within their interval of convergence, power series can be handled much like polynomials. (1) A power series converges uniformly and absolutely in any interval which lies entirely within its interval of convergence. (2) A power series can be di/C128erentiated or integrated term by term over any interval lying entirely within the interval of convergence. Also, the sum of a convergent power series is continuous in any interval lying entirely within its interval of convergence. (3) Two power series can be added or subtracted term by term for each value of xcommon to their intervals of convergence. (4) Two power series, for example,P1 nˆ0anxnandP1nˆ0bnxn, can be multiplied to obtainP1 nˆ0cnxn;where cnˆa0bn‡a1bnÿ1‡a2bnÿ2‡‡ anb0, the result being, valid for each xwithin the common interval of convergence. (5) If the power seriesP1nˆ0anxnis divided by the power seriesP1nˆ0bnxn, where b06ˆ0, the quotient can be written as a power series which converges for suciently small values of x. /C84a/C121lor/C39s e/C120pansion It is very useful in most applied work to find power series that represent the given functions. We now review one method of obtaining such series, the Taylor expan- sion. We assume that our function f…x†has a continuous nth derivative in the interval /C91 a,b/C93 and that there is a Taylor series for f…x†of the form f…x†ˆa0‡a1…xÿ †‡a2…xÿ †2‡a3…xÿ †3‡‡ an…xÿ †n‡ ; …A1:8† where lies in the interval /C91 a,b/C93. Di/C128erentiating, we have f0…x†ˆa1‡2a2…xÿ †‡3a3…xÿ †2‡‡ nan…xÿ †nÿ1‡ ; f00…x†ˆ2a2‡32a3…xÿ †‡43a4…xÿa†2‡‡ n…nÿ1†an…xÿ †nÿ2‡ ; ... f…n†…x†ˆn…nÿ1†…nÿ2†1an‡terms containing powers of …xÿ †: We now put xˆ in each of the above derivatives and obtain f… †ˆa0;f0… †ˆa1;f00… †ˆ2a2;fF… †ˆ3/C33a3;;f…n†… †ˆn/C33an; 524APPENDI/C88 1 PRELIMINARIES where f0… †means that f…x†has been di/C128erentiated and then we have put xˆ ; and by f00… †we mean that we have found f00…x†and then put xˆ , and so on. Substituting these into (A1.8) we obtain f…x†ˆf… †‡f0… †…xÿ †‡1 2/C33f00… †…xÿ †2‡‡1 n/C33f…n†… †…xÿ †n‡ : …A1:9† This is the Taylor series for f…x†about xˆ . The Maclaurin series for f…x†is the Taylor series about the origin. Putting ˆ0 in (A1.9), we obtain the Maclaurin series for f…x†: f…x†ˆf…0†‡f0…0†x‡1 2/C33f00…0†x2‡1 3/C33fF…0†x3‡‡1 n/C33f…n†…0†xn‡ : …A1:10† Example A1.15 Find the Maclaurin series expansion of the exponential function ex. Solution: Here f…x†ˆex. Di/C128erentiating, we have f…n†…0†ˆ1 for all n;nˆ1;2;3.... Then, by Eq. (A1.10), we have exˆ1‡x‡1 2/C33x2‡1 3/C33x3‡ ˆX1 nˆ0xn n/C33; ÿ1 <x<1: The following series are frequently employed in practice: …1†sinxˆxÿx3 3/C33‡x5 5/C33ÿx7 7/C33‡… ÿ 1†nÿ1x2nÿ1 …2nÿ1†/C33‡ ; ÿ1 <x<1: …2†cosxˆ1ÿx2 2/C33‡x4 4/C33ÿx6 6/C33‡… ÿ 1†nÿ1x2nÿ2 …2nÿ2†/C33‡ ; ÿ1 <x<1: …3†exˆ1‡x‡x2 2/C33‡x3 3/C33‡‡xnÿ1 …nÿ1†/C33‡ ; ÿ1 <x<1: …4†lnj…1‡xjˆxÿx2 2‡x3 3ÿx4 4‡… ÿ 1†nÿ1xn n‡ ; ÿ1<x1: …5†1 2ln1‡x 1ÿx/C12/C12/C12/C12/C12/C12/C12/C12ˆx‡x3 3‡x5 5‡x7 7‡‡x2nÿ1 2nÿ1‡ ; ÿ1<x<1: …6†tanÿ1xˆxÿx3 3‡x5 5ÿx7 7‡… ÿ 1†nÿ1x2nÿ1 2nÿ1‡ ; ÿ1x1: …7†…1‡x†/C112ˆ1‡/C112x‡/C112…/C112ÿ1† 2/C33x2‡‡/C112…/C112ÿ1†… /C112ÿn‡1† n/C33xn‡ : 525TAYLOR’S E/C88PANSION This is the binomial series: ( a)I fpis a positive integer or zero, the series termi- nates. ( b)I f/C112/C620 but is not an integer, the series converges absolutely for ÿ1x1. (c)I f ÿ1</C112<0, the series converges for ÿ1<x1. (d)I f /C112ÿ1, the series converges for ÿ1<x<1. Problem A1.16 Obtain the Maclaurin series for sin x(the Taylor series for sin xabout xˆ0). Problem A1.17Use series methods to obtain the approximate value ofR 1 0…1ÿeÿx†=xdx. We can find the power series of functions other than the most common ones listed above by the successive di/C128erentiation process given by Eq. (A1.9). Thereare simpler ways to obtain series expansions. We give several useful methods here. (a) For example to find the series for …x‡1†sinx, we can multiply the series for sinxby…x‡1†and collect terms: …x‡1†sinxˆ…x‡1†xÿx3 3/C33‡x5 5/C33ÿ/C32! ˆx‡x2ÿx3 3/C33ÿx4 3/C33‡ : To find the expansion for excosx, we can multiply the series for exby the series for cos x: excosxˆ1‡x‡x2 2/C33‡x3 3/C33‡/C32! 1ÿx2 2/C33‡x4 4/C33‡/C32! ˆ1‡x‡x2 2/C33‡x3 3/C33‡x4 4/C33 ÿx2 2/C33ÿx3 3/C33ÿx4 2/C332/C33 ‡x4 4/C33 ˆ1‡xÿx3 3ÿx4 6: Note that in the first example we obtained the desired series by multiplication of a known series by a polynomial; and in the second example we obtained thedesired series by multiplication of two series. (b) In some cases, we can find the series by division of two series. For example, to find the series for tan x, we can divide the series for sinx by the series for cos x: 526APPENDI/C88 1 PRELIMINARIES tanxˆsinx cosxˆ 1ÿx2 2‡x4 4/C33/C30 xÿx3 3/C33‡x5 5/C33 ˆx‡1 3x3‡2 15x5: The last step is by long division x‡13x 3‡2 15x5 1ÿx2 2/C33‡x4 4/C33 xÿx3 3/C33‡x5 5/C33/C115 xÿx3 2/C33‡x5 4/C33 x3 3ÿx5 30 x3 3ÿx5 6 2x5 15;etc: Problem A1.18 Find the series expansion for 1 =…1‡x†by long division. Note that the series can be found by using the binomial series: 1 =…1‡x†ˆ… 1‡x†ÿ1. (c) In some cases, we can obtain a series expansion by substitution of a poly- nomial or a series for the variable in another series. As an example, let us find the series for eÿx2. We can replace xin the series for exbyÿx2and obtain eÿx2ˆ1ÿx2‡…ÿx2†2 2/C33‡…ÿx†3 3/C33ˆ 1ÿx2‡x4 2/C33ÿx6 3/C33: Similarly, to find the series for sinxp=xpwe replace xin the series for sin xbyxpand obtain sinxp xp ˆ1ÿx 3/C33‡x2 5/C33; x/C620: Problem A1.19 Find the series expansion for etanx. Problem A1.20 Assuming the power series for exholds for complex numbers, show that eixˆcosx‡isinx: 527TAYLOR’S E/C88PANSION (d) Find the series for tanÿ1x(arc tan x). We can find the series by the succes- sive di/C128erentiation process. But it is very tedious to find successive derivatives of tanÿ1x. We can take advantage of the following integration Zx 0dt 1‡t2ˆtanÿ1t/C12/C12/C12/C12x 0ˆtanÿ1x: We now first write out …1‡t2†ÿ1as a binomial series and then integrate term by term: Zx 0dt 1‡t2ˆZx 01ÿt2‡t4ÿt6‡ÿ dtˆtÿt3 3‡t5 5ÿt7 7‡/C12/C12/C12/C12x 0: Thus, we have tanÿ1xˆxÿx3 3‡x5 5ÿx7 7‡ : (e) Find the series for ln xabout xˆ1. We want a series of powers ( xÿ1) rather than powers of x. We first write lnxˆln‰1‡…xÿ1†Š and then use the series ln …1‡x†with xreplaced by ( xÿ1): lnxˆln‰1‡…xÿ1†Š ˆ … xÿ1†ÿ1 2…xÿ1†2‡13…xÿ1† 3ÿ14…xÿ1† 4: Problem A1.21 Expand cos xabout xˆ3=2. /C72igher deri/C118ati/C118es and Leibnit/C122/C39s formula for nth deri/C118ati/C118e of a product Higher derivatives of a function yˆf…x†with respect to xare written as d2y dx2ˆd dxdy dx ;d3y dx3ˆd dxd2y dx2/C32! ; ...;dny dxnˆd dydnÿ1y dxnÿ1/C32! : These are sometimes abbreviated to either f00…x†;fF…x†;...;f…n†…x†orD2y;D3y;...;Dny where Dˆd=dx. When higher derivatives of a product of two functions f…x†and /C103…x†are required, we can proceed as follows: D…f/C103†ˆfD/C103‡/C103Df 528APPENDI/C88 1 PRELIMINARIES and D2…f/C103†ˆD…fD/C103‡/C103Df†ˆfD2/C103‡2DfD/C103‡D2/C103: Similarly we obtain D3…f/C103†ˆfD3/C103‡3DfD2/C103‡3D2fD/C103‡/C103D3f; D4…f/C103†ˆfD4/C103‡4DfD3/C103‡6D2fD2/C103‡4D3D/C103‡/C103D4/C103; and so on. By inspection of these results the following formula (due to Leibnitz) may be written down for nth derivative of the product fg: Dn…f/C103†ˆf…Dn/C103†‡n…Df†…Dnÿ1/C103†‡n…nÿ1† 2/C33…D2f†…Dnÿ2/C103†‡ ‡n/C33 k/C33…nÿk†/C33…Dkf†…Dnÿkg†‡‡… Dnf†/C103: Example A1.16 Iffˆ1ÿx2;/C103ˆD2y, where yis a function of x, say u…x†, then Dnf…1ÿx2†D2ygˆ… 1ÿx2†Dn‡2yÿ2nxDn‡1yÿn…nÿ1†Dny: Leibnitz’s formula may also be applied to a di/C128erential equation. For example, ysatisfies the di/C128erential equation D2y‡x2yˆsinx: Then di/C128erentiating each term ntimes we obtain Dn‡2y‡…x2Dny‡2nxDnÿ1y‡n…nÿ1†Dnÿ2y†ˆsinn 2‡x ; where we have used Leibnitz’s formula for the product term x2y. Problem A1.22Using Leibnitz’s formula, show that D n…x2sinx†ˆf x2ÿn…nÿ1†gsin…x‡n=2†ÿ2nxcos…x‡n=2†: /C83ome important properties of definite integrals Integration is an operation inverse to that of di/C128erentiation; and it is a device forcalculating the ‘area under a curve’. The latter method regards the integral as the limit of a sum and is due to Riemann. We now list some useful properties of definite integrals. (1) If in axb;mf…x†M, where mandMare constants, then m…bÿa†Z b af…x†dM…bÿa†: 529PROPERTIES OF DEFINITE INTEGRALS Divide the interval /C91 a,b/C93 into nsubintervals by means of the points x1;x2;...;xnÿ1chosen arbitrarily. Let /C17kbe any point in the subinterval xkÿ1/C17kxk, then we have mxkf…/C17k†xkMxk;kˆ1;2;...;n; where xkˆxkÿxkÿ1. Summing from kˆ1t o nand using the fact that Xn kˆ1xkˆ…x1ÿa†‡…x2ÿx1†‡‡… bÿxnÿ1†ˆbÿa; it follows that m…bÿa†Xn rkˆ1f…/C17k†xkM…bÿa†: Taking the limit as n!1 and each xk!0 we have the required result. (2) If in axb;f…x†/C103…x†, then Zb af…x†dxZb a/C103…x†dx: (3)Zb af…x†dx/C12/C12/C12/C12/C12/C12/C12/C12Z b af…x†jj dx ifa<b: From the inequality a‡b‡c‡ jj ajj‡bjj‡cjj‡ ; where jajis the absolute value of a real number a, we have Xn kˆ1f…/C17k†xk/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12X n kˆ1f…/C17k†xk jj ˆXn kˆ1f…/C17k†jj xk: Taking the limit as n!1 and each xk!0 we have the required result. (4) The mean value theorem: If f…x†is continuous in /C91 a,b/C93, we can find a point /C17 in (a,b) such that Zb af…x†dxˆ…bÿa†f…/C17†: Since f…x†is continuous in /C91 a,b/C93, we can find constants mandMsuch that mf…x†M. Then by (1) we have m1 bÿaZb af…x†dxM: Since f…x†is continuous it takes on all values between mandM; in parti- cular there must be a value /C17such that f…/C17†ˆZb af…x†dx=…bÿa†;a</C17< b: The required result follows on multiplying by bÿa. 530APPENDI/C88 1 PRELIMINARIES /C83ome useful methods of integration (1) Changing variables: We use a simple example to illustrate this common procedure. Consider the integral IˆZ1 0eÿax2dx; which is equal to …=a†1=2=2:To show this let us write IˆZ1 0eÿax2dxˆZ1 0eÿay2dy: Then I2ˆZ1 0eÿax2dxZ1 0eÿay2dyˆZ1 0Z1 0eÿa…x2‡y2†dxdy : We now rewrite the integral in plane polar coordinates …r;†/C58 x2‡y2ˆr2;dxdy ˆrdrd. Then I2ˆZ1 0Z=2 0eÿar2rddrˆ 2Z1 0eÿar2rdrˆ 2 ÿeÿar2 2a/C12/C12/C12/C121 0ˆ 4a and IˆZ1 0eÿax2dxˆ…=a†1=2=2: (2) Integration by parts: Since d dxu/C118…† ˆ ud/C118 dx‡/C118du dx; where uˆf…x†and/C118ˆ/C103…x†, it follows that Z ud/C118 dx dxˆu/C118ÿZ /C118du dx dx: This can be a useful formula in evaluating integrals. Example A1.17 Evaluate IˆR tanÿ1xdx Solution: Since tanÿ1x can be easily di/C128erentiated, we write IˆR tanÿ1xdxˆR 1tanÿ1xdxand let uˆtanÿ1x;dv=dxˆ1. Then Iˆxtanÿ1xÿZxdx 1‡x2ˆxtanÿ1xÿ1 2log…1‡x2†‡c: 531SOME USEFUL METHODS OF INTEGRATION Example A1.18 Show that Z1 ÿ1x2eÿax2dxˆ1=2 2a3=2: Solution: Let us first consider the integral IˆZ1 0eÿax2dx ‘Integration-by-parts’ gives IˆZc beÿax2dxˆeÿax2x/C12/C12/C12/C12c b‡2Zc bax2eÿax2dx; from which we obtain Zc bx2eÿax2dxˆ1 2aZc beÿax2dxÿeÿax2x/C12/C12/C12/C12c b : We let limits bandcbecome ÿ1 and‡1, and thus obtain the desired result. Problem A1.23 Evaluate IˆZ xexdx(constant). (3) Partial fractions: Any rational function P…x†=/C81…x†, where P…x†and/C81…x† are polynomials, with the degree of P…x†less than that of /C81…x†, can be written as the sum of rational functions having the form A=…ax‡b†k, …Ax‡B†=…ax2‡bx‡c†k, where kˆ1;2;3;...which can be integrated in terms of elementary functions. Example A1.19 3xÿ2 …4xÿ3†…2x‡5†3ˆA 4xÿ3‡B …2x‡5†3‡C …2x‡5†2‡D 2x‡5; 5x2ÿx‡2 …x2‡2x‡4†2…xÿ1†ˆAx‡B …x2‡2x‡4†2‡Cx‡D x2‡2x‡4‡/C69 xÿ1: Solution: The coecients A,B,/C67etc., can be determined by clearing the frac- tions and equating coecients of like powers of xon both sides of the equation. Problem A1.24 Evaluate IˆZ6ÿx …xÿ3†…2x‡5†dx: 532APPENDI/C88 1 PRELIMINARIES (4) Rational functions of sin xand cos xcan always be integrated in terms of elementary functions by substitution tan …x=2†ˆu, as shown in the following example. Example A1.20 Evaluate IˆZdx 5‡3c o s x: Solution: Let tan …x=2†ˆu, then sin…x=2†ˆu= 1‡u2/C112 ;cos…x=2†ˆ1=1‡u 2/C112 and cosxˆcos2…x=2†ÿsin2…x=2†ˆ1ÿu2 1‡u2; also duˆ1 2sec2…x=2†dx ordxˆ2 cos2…x=2†ˆ2du=…1‡u2†: Thus IˆZdu u2‡4ˆ1 2tanÿ1…u=2†‡cˆ12tanÿ1‰12tanx=2†Š ‡c: /C82eduction formulas Consider an integral of the formR xneÿxdx. Since this depends upon nlet us call it In. Then using integration by parts we have Inˆÿxneÿx‡nZ xnÿ1eÿxdxˆxneÿx‡nInÿ1: The above equation gives Inin terms of Inÿ1…Inÿ2;Inÿ3, etc.) and is therefore called a reduction formula. Problem A1.25 Evaluate InˆZ=2 0sinnxdxˆZ=2 0sinxsinnÿ1xdx: 533REDUCTION FORMULAS /C68i/C128erentiation of integrals (1) Indefinite integrals: We first consider di/C128erentiation of indefinite integrals. If f…x; †is an integrable function of xand is a variable parameter, and if Z f…x; †dxˆ/C71…x; †; …A1:11† then we have /C64/C71…x; †=/C64xˆf…x; †: …A1:12† Furthermore, if f…x; †is such that /C642/C71…x; † /C64x/C64 ˆ/C642/C71…x; † /C64 /C64x; then we obtain /C64 /C64x/C64/C71…x; † /C64  ˆ/C64 /C64 /C64/C71…x; † /C64x ˆ/C64f…x; † /C64 and integrating gives Z/C64f…x; † /C64 dxˆ/C64/C71…x; † /C64 ; …A1:13† which is valid provided /C64f…x; †=/C64 is continuous in xas well as . (2) Definite integrals: We now extend the above procedure to definite integrals: I… †ˆZb af…x; †dx; …A1:14† where f…x; †is an integrable function of xin the interval axb,a n d aandb are in general continuous and di/C128erentiable (at least once) functions of . We now have a relation similar to Eq. (A1.11): I… †ˆZb af…x; †dxˆ/C71…b; †ÿ/C71…a; †… A1:15† and, from Eq. (A1.13), Zb a/C64f…x; † /C64 dxˆ/C64/C71…b; † /C64 ÿ/C64/C71…a; † /C64 : …A1:16† Di/C128erentiating (A1.15) totally dI… † d ˆ/C64/C71…b; † /C64bdb d ‡/C64/C71…b; † /C64 ÿ/C64/C71…a; † /C64ada d ÿ/C64/C71…a; † /C64 /C58 534APPENDI/C88 1 PRELIMINARIES which becomes, with the help of Eqs. (A1.12) and (A1.16), dI… † d ˆZb a/C64f…x; † /C64 dx‡f…b; †db d ÿf…a; †da d ; …A1:17† which is known as Leibnitz’s rule for di/C128erentiating a definite integral. If aandb, the limits of integration, do not depend on , then Eq. (A1.17) reduces to dI… † d ˆd d Zb af…x; †dxˆZb a/C64f…x; † /C64 dx: Problem A1.26 IfI… †ˆZ 2 0sin… x†=xdx, find dI=d . /C72omogeneous functions A homogeneous function f…x1;x2;...;xn†of the kth degree is defined by the relation f…x1;x2;...;xn†ˆkf…x1;x2;...;xn†: For example, x3‡3x2yÿy3is homogeneous of the third degree in the variables x andy. Iff…x1;x2;...;xn) is homogeneous of degree kthen it is straightforward to show that Xn jˆ1xj/C64f /C64xjˆkf: This is known as Euler’s theorem on homogeneous functions. Problem A1.27 Show that Euler’s theorem on homogeneous functions is true. /C84a/C121lor series for functions of t/C119o independent /C118ariables The ideas involved in Taylor series for functions of one variable can be general-ized. For example, consider a function of two variables ( x;y). If all the nth partial derivatives of f…x;y†are continuous in a closed region and if the …n‡1)st partial 535HOMOGENEOUS FUNCTIONS derivatives exist in the open region, then we can expand the function f…x;y†about xˆx0;yˆy0in the form f…x0‡/C104;y0‡k†ˆf…x0;y0†‡ /C104/C64 /C64x‡k/C64 /C64y f…x0;y0† ‡1 2/C33/C104/C64 /C64x‡k/C64 /C64y2 f…x0;y0† ‡‡1 n/C33/C104/C64 /C64x‡k/C64 /C64yn f…x0;y0†‡Rn; where /C104ˆxˆxÿx0;kˆyˆyÿy0;Rn, the remainder after nterms, is given by Rnˆ1 …n‡1†/C33/C104/C64 /C64x‡k/C64 /C64yn‡1 f…x0‡/C104;y0‡k†;0<< 1; and where we use the operator notation /C104/C64 /C64x‡k/C64 /C64y f…x0;y0†ˆ/C104fx…x0;y0†‡kfy…x0;y0†; /C104/C64 /C64x‡k/C64 /C64y2 f…x0;y0†ˆ /C1042/C642 /C64x2‡2/C104k/C642 /C64x/C64y‡k2/C642 /C64y2/C32! f…x0;y0†; etc., when we expand /C104/C64 /C64x‡k/C64 /C64yn formally by the binomial theorem. When lim n!1Rnˆ0 for all …x;y) in a region, the infinite series expansion is called a Taylor series in two variables. Extensions can be made to three or more variables. Lagrange multiplier For functions of one variable such as f…x†to have a stationary value (maximum or minimum) at xˆa, we have f0…a†ˆ0. If fn…a†<0 it is a relative maximum while if f…a†/C620 it is a relative minimum. Similarly f…x;y†has a relative maximum or minimum at xˆa;yˆbif fx…a;b†ˆ0,fy…a;b†ˆ0. Thus possible points at which f…x;y†has a relative max- imum or minimum are obtained by solving simultaneously the equations /C64f=/C64xˆ0;/C64 f=/C64yˆ0: 536APPENDI/C88 1 PRELIMINARIES Sometimes we wish to find the relative maxima or minima of f…x;y†ˆ0 subject to some constraint condition /C30…x;y†ˆ0. To do this we first form the function /C103…x;y†ˆf…x;y†‡f…x;y†and then set /C64/C103=/C64xˆ0;/C64 /C103=/C64yˆ0: The constant is called a Lagrange multiplier and the method is known as the method of undetermined multipliers. 537LAGRANGE MULTIPLIER Appendix 2 /C68eterminants The determinant is a tool used in many branches of mathematics, science, and engineering. The reader is assumed to be familiar with this subject. However, for those who are in need of review, we prepared this appendix, in which the deter- minant is defined and its properties developed. In Chapters 1 and 3, the reader will see the determinant’s use in proving certain properties of vector and matrix operations. The concept of a determinant is already familiar to us from elementary algebra, where, in solving systems of simultaneous linear equation, we find it convenient to use determinants. For example, consider the system of two simultaneous linear equations a11x1‡a12x2ˆb1; a21x1‡a22x2ˆb2;) …A2:1† in two unknowns x1;x2where aij…i;jˆ1;2†are constants. These two equations represent two lines in the x1x2plane. To solve the system (A2.1), multiplying the first equation by a22, the second by ÿa12and then adding, we find x1ˆb1a22ÿb2a12 a11a22ÿa21a12: …A2:2a† Next, by multiplying the first equation by ÿa21, the second by a11and adding, we find x2ˆb2a11ÿb1a21 a11a22ÿa21a12: …A2:2b† We may write the solutions (A2.2) of the system (A2.1) in the determinant form x1ˆD1 D;x2ˆD2 D; …A2:3† 538 where D1ˆb1a12 b2a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;D 2ˆa11b1 a21b2/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;Dˆa 11a12 a21a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12…A2:4† are called determinants of second order or order 2. The numbers enclosed between vertical bars are called the elements of the determinant. The elements in ahorizontal line form a row and the elements in a vertical line form a column of the determinant. It is obvious that in Eq. (A2.3) D6ˆ0. Note that the elements of determinant /C68are arranged in the same order as they occur as coecients in Eqs. (A1.1). The numerator D 1forx1is constructed from /C68by replacing its first column with the coecients b1andb2on the right-hand side of (A2.1). Similarly, the numerator for x2is formed by replacing the second column of /C68byb1;b2. This procedure is often called Cramer’s rule. Comparing Eqs. (A2.3) and (A2.4) with Eq. (A2.2), we see that the determinant is computed by summing the products on the rightward arrows and subtractingthe products on the leftward arrows: a 11a12 a21a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆa 11a22ÿa12a21;etc: …ÿ† …‡† This idea is easily extended. For example, consider the system of three linear equations a11x1‡a12x2‡a13x3ˆb1; a21x1‡a22x2‡a23x3ˆb2; a31x1‡a32x2‡a33x3ˆb3;9 >>= >>;…A2:5† in three unknowns x1;x2;x3. To solve for x1, we multiply the equations by a22a33ÿa32a23;ÿ…a12a33ÿa32a13†;a12a23ÿa22a13; respectively, and then add, finding x1ˆb1a22a33ÿb1a23a32‡b2a13a32ÿb2a12a33‡b3a12a23ÿb3a13a22 a11a22a33ÿa11a32a23‡a21a32a13ÿa21a12a33‡a31a12a23ÿa31a22a13; which can be written in determinant form x1ˆD1=D; …A2:6† 539APPENDI/C88 2 DETERMINANTS where Dˆa11a12a13 a21a22a23 a31a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;D 1ˆb1a12a13 b2a22a23 b3a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: …A2:7† Again, the elements of /C68are arranged in the same order as they appear as coecients in Eqs. (A2.5), and D 1is obtained by Cramer’s rule. In the same manner we can find solutions for x2;x3. Moreover, the expansion of a determi- nant of third order can be obtained by diagonal multiplication by repeating on the right the first two columns of the determinant and adding the signed products of the elements on the various diagonals in the resulting array: This method of writing out determinants is correct only for second- and third- order determinants. Problem A2.1 Solve the following system of three linear equations using Cramer’s rule: 2x1ÿx2‡2x3ˆ2; x1‡10x2ÿ3x3ˆ5; ÿx1‡x2‡x3ˆÿ3: Problem A2.2 Evaluate the following determinants …a†12 43/C12/C12/C12/C12/C12/C12/C12/C12;…b†51 8 1 536 1 042/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;…c†cos ÿsin sin cos /C12/C12/C12/C12/C12/C12/C12/C12: /C68eterminants/C44 minors/C44 and cofactors We are now in a position to define an nth-order determinant. A determinant of order nis a square array of n 2quantities enclosed between vertical bars, 540APPENDI/C88 2 DETERMINANTS a11a12a13 a21a22a23 a31a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12a 11 a21 a31a12 a22 a32 …ÿ† …ÿ† …ÿ† …‡† …‡† …‡† Dˆa11a12 a1n a21a22 a2n ......... an1an2 ann/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: …A2:8† By deleting the ith row and the kth column from the determinant /C68we obtain an (nÿ1)st order determinant (a square array of nÿ1 rows and nÿ1 columns between vertical bars), which is called the minor of the element a ik(which belongs to the deleted row and column) and is denoted by Mik. The minor Mikmultiplied by…ÿ†i‡kis called the cofactor of aikand is denoted by Cik: Cikˆ… ÿ 1†i‡kMik: …A2:9† For example, in the determinant a11a12a13 a21a22a23 a31a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; we have C 11ˆ… ÿ 1†1‡1M11ˆa22a23 a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; C 32ˆ… ÿ 1†3‡2M32ˆÿa11a13 a21a23/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;etc: It is very convenient to get the proper sign (plus or minus) for the cofactor …ÿ1† i‡kby thinking of a checkerboard of plus and minus signs like this ‡ÿ‡ ÿ ÿ‡ÿ ‡ ‡ÿ‡ ÿ etc: ÿ‡ÿ ‡ etc:... ‡ÿ ÿ‡/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 thus, for the element a 23we can see that the checkerboard sign is minus. /C69/C120pansion of determinants Now we can see how to find the value of a determinant: multiply each of one row (or one column) by its cofactor and then add the results, that is, 541E/C88PANSION OF DETERMINANTS Dˆai1Ci1‡ai2Ci2‡  ‡ ainCin ˆXn kˆ1aikCik …iˆ1;2;...;orn†… A2:10a† …cofactor expansion along the ith row † or Dˆa1kC1k‡a2kC2k‡‡ ankCnk ˆXn iˆ1aikCik …kˆ1;2;...;orn†: …A2:10b† …cofactor expansion along the kth column † We see that /C68is defined in terms of ndeterminants of order nÿ1, each of which, in turn, is defined in terms of nÿ1 determinants of order nÿ2, and so on; we finally arrive at second-order determinants, in which the cofactors of the elements are single elements of /C68. The method of evaluating a determinant just described is one form of Laplace’s development of a determinant. Problem A2.3 For a second-order determinant Dˆa11a12 a21a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12 show that the Laplace’s development yields the same value of /C68no matter which row or column we choose. Problem A2.4 Let Dˆ130 264 ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: Evaluate /C68, first by the first-row expansion, then by the first-column expansion. Do you get the same value of /C68/C63 Properties of determinants In this section we develop some of the fundamental properties of the determinant function. In most cases, the proofs are brief. 542APPENDI/C88 2 DETERMINANTS (1) If all elements of a row (or a column) of a determinant are zero, the value of the determinant is zero. Proof: Let the elements of the kth row of the determinant /C68be zero. If we expand /C68in terms of the ith row, then Dˆai1Ci1‡ai2Ci2‡  ‡ ainCin: Since the elements ai1;ai2;...;ainare zero, Dˆ0. Similarly, if all the elements in one column are zero, expanding in terms of that column shows that the determi- nant is zero. (2) If all the elements of one row (or one column) of a determinant are multi- plied by the same factor k, the value of the new determinant is ktimes the value of the original determinant. That is, if a determinant Bis obtained from determinant /C68by multiplying the elements of a row (or a column) of /C68by the same factor k, then BˆkD. Proof: Suppose Bis obtained from /C68by multiplying its ith row by k. Hence the ith row of Biskaij, where jˆ1;2;...;n, and all other elements of Bare the same as the corresponding elements of A. Now expand Bin terms of the ith row: Bˆkai1Ci1‡kai2Ci2‡  ‡ kainCin ˆk…ai1Ci1‡ai2Ci2‡  ‡ ainCin† ˆkD: The proof for columns is similar. Note that property (1) can be considered as a special case of property (2) with kˆ0. Example A2.1 If Dˆ12 3 01 1 4ÿ10/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12and Bˆ16 3 03 1 4ÿ30/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; then we see that the second column of Bis three times the second column of /C68. Evaluating the determinants, we find that the value of /C68isÿ3, and the value of B isÿ9 which is three times the value of /C68, illustrating property (2). Property (2) can be used for simplifying a given determinant, as shown in the following example. 543PROPERTIES OF DETERMINANTS Example A2.2 130 264 ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ2130 132 ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ23110 112 ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ232110 111 ÿ101/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆÿ12: (3) The value of a determinant is not altered if its rows are written as columns, in the same order. Proof: Since the same value is obtained whether we expand a determinant by any row or any column, thus we have property (3). The following example will illustrate this property. Example A2.3 Dˆ10 2 ÿ11 0 2ÿ13/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ110 ÿ13/C12/C12/C12/C12/C12/C12/C12/C12ÿ0ÿ10 23/C12/C12/C12/C12/C12/C12/C12/C12‡2ÿ11 2ÿ1/C12/C12/C12/C12/C12/C12/C12/C12ˆ1: Now interchanging the rows and the columns, then evaluating the value of the resulting determinant, we find 1ÿ12 01 ÿ1 203/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ11ÿ1 03/C12/C12/C12/C12/C12/C12/C12/C12ÿ… ÿ 1†0ÿ1 23/C12/C12/C12/C12/C12/C12/C12/C12‡201 20/C12/C12/C12/C12/C12/C12/C12/C12ˆ1; illustrating property (3). (4) If any two rows (or two columns) of a determinant are interchanged, the resulting determinant is the negative of the original determinant. Proof: The proof is by induction. It is easy to see that it holds for 2 2 deter- minants. Assuming the result holds for nndeterminants, we shall show that it also holds for …n‡1†… n‡1†determinants, thereby proving by induction that it holds in general. LetBbe an …n‡1†…n‡1†determinant obtained from /C68by interchanging two rows. Expanding B in terms of a row that is not one of those interchanged, such as the kth row, we have BˆX n jˆ1…ÿ1†j‡kbkjM0 kj; where M0 kjis the minor of bkj. Each bkjis identical to the corresponding akj(the elements of /C68). Each M0 kjis obtained from the corresponding Mkj(ofakj)b y 544APPENDI/C88 2 DETERMINANTS interchanging two rows. Thus bkjˆakj, and M0 kjˆÿMkj. Hence BˆÿXn jˆ1…ÿ1†j‡kbkjMkjˆÿD: The proof for columns is similar. Example A2.4 Consider Dˆ10 2 ÿ11 0 2ÿ13/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ1: Now interchanging the first two rows, we have Bˆÿ11 0 10 2 2ÿ13/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆÿ1 illustrating property (4). (5) If corresponding elements of two rows (or two columns) of a determinant are proportional, the value of the determinant is zero. Proof: Let the elements of the ith and jth rows of /C68be proportional, say, a ikˆcajk;kˆ1;2;...;n.I fcˆ0, then /C68ˆ0. For c6ˆ0, then by property (2), /C68ˆcB, where the ith and jth rows of Bare identical. Interchanging these two rows, Bgoes over to ÿB(by property (4)). But the rows are identical, the new determinant is still B. Thus BˆÿB;Bˆ0, and /C68ˆ0. Example A2.5 Bˆ11 2 ÿ1ÿ10 22 8/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ0;Dˆ36 ÿ4 1ÿ13 ÿ6ÿ12 8/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ0: InBthe first and second columns are identical, and in /C68the first and the third rows are proportional. (6) If each element of a row of a determinant is a binomial, then the determi- nant can be written as the sum of two determinants, for example, 4x‡232 x 43 3xÿ121/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆ4x32 x43 3x21/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12‡232 043 ÿ121/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: 545PROPERTIES OF DETERMINANTS Proof: Expanding the determinant by the row whose terms are binomials, we will see property (6) immediately. (7) If we add to the elements of a row (or column) any constant multiple of the corresponding elements in any other row (or column), the value of the determi- nant is unaltered. Proof: Applying property (6) to the determinant that results from the given addition, we obtain a sum of two determinants: one is the original determinant and the other contains two proportional rows. Then by property (4), the seconddeterminant is zero, and the proof is complete. It is advisable to simplify a determinant before evaluating it. This may be done with the help of properties (7) and (2), as shown in the following example. Example A2.6 Evaluate Dˆ12 4 2 1 9 3 2ÿ37ÿ1 194 ÿ23 50 ÿ171 ÿ3 177 63 234/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: To simplify this, we want the first elements of the second, third and last rows all to be zero. To achieve this, add the second row to the third, and add three times thefirst to the last, subtract twice the first row from the second; then develop the resulting determinant by the first column: Dˆ12 42 1 9 3 0ÿ85 ÿ43 8 0ÿ2 ÿ12 3 0 249 126 513/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆÿ85ÿ43 8 ÿ2ÿ12 3 249 126 513/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: We can simplify the resulting determinant further. Add three times the first row to the last row: Dˆÿ85ÿ43 8 ÿ2ÿ12 3 ÿ6ÿ3 537/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: Subtract twice the second column from the first, and then develop the resulting determinant by the first column: Dˆ1ÿ43 8 0ÿ12 3 0ÿ3 537/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆÿ12 3 ÿ3 537/C12/C12/C12/C12/C12/C12/C12/C12ˆÿ537ÿ23… ÿ 3†ˆÿ 468 : 546APPENDI/C88 2 DETERMINANTS By applying the product rule of di/C128erentiation we obtain the following theorem. /C68eri/C118ati/C118e of a determinant If the elements of a determinant are di/C128erentiable functions of a variable, then the derivative of the determinant may be written as a sum of individual determinants, for example, d dxabc ef /C103 /C104mn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ˆa 0b0c0 ef/C103 /C104mn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12‡abc e 0f0/C1030 /C104mn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12‡abc ef /C103 /C104 0m0n0/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; where a;b;...;m;nare di/C128erentiable functions of x, and the primes denote deri- vatives with respect to x. Problem A2.5 Show, without computation, that the following determinants are equal to zero: 0 aÿb ÿa 0 c bÿc 0/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;02 ÿ3 ÿ204 3ÿ40/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12: Problem A2.6 Find the equation of a plane which passes through the three points (0, 0, 0), (1, 2, 5), and (2, ÿ1, 0). 547DERIVATIVE OF A DETERMINANT Appendix 3 Table of /C42 /C70…x†ˆ1  2pZx 0eÿt2=2dt: 548x 0.0 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09 0.0 0.0000 0.0040 0.0080 0.0120 0.0160 0.0199 0.0239 0.0279 0.0319 0.0359 0.1 0.0398 0.0438 0.0478 0.0517 0.0557 0.0596 0.0636 0.0675 0.0714 0.07530.2 0.0793 0.0832 0.0871 0.0910 0.0948 0.0987 0.1026 0.1064 0.1103 0.11410.3 0.1179 0.1217 0.1255 0.1293 0.1331 0.1368 0.1406 0.1443 0.1480 0.15170.4 0.1554 0.1591 0.1628 0.1664 0.1700 0.1736 0.1772 0.1808 0.1844 0.18790.5 0.1915 0.1950 0.1985 0.2019 0.2054 0.2088 0.2123 0.2157 0.2190 0.2224 0.6 0.2257 0.2291 0.2324 0.2357 0.2389 0.2422 0.2454 0.2486 0.2517 0.2549 0.7 0.2580 0.2611 0.2642 0.2673 0.2704 0.2734 0.2764 0.2794 0.2823 0.28520.8 0.2881 0.2910 0.2939 0.2967 0.2995 0.3023 0.3051 0.3078 0.3106 0.31330.9 0.3159 0.3186 0.3212 0.3238 0.3264 0.3289 0.3315 0.3340 0.3365 0.3389 1.0 0.3413 0.3438 0.3461 0.3485 0.3508 0.3531 0.3554 0.3577 0.3599 0.3621 1.1 0.3643 0.3665 0.3686 0.3708 0.3729 0.3749 0.3770 0.3790 0.3810 0.3830 1.2 0.3849 0.3869 0.3888 0.3907 0.3925 0.3944 0.3962 0.3980 0.3997 0.40151.3 0.4032 0.4049 0.4066 0.4082 0.4099 0.4115 0.4131 0.4147 0.4162 0.41771.4 0.4192 0.4207 0.4222 0.4236 0.4251 0.4265 0.4279 0.4292 0.4306 0.43191.5 0.4332 0.4345 0.4357 0.4370 0.4382 0.4394 0.4406 0.4418 0.4429 0.4441 1.6 0.4452 0.4463 0.4474 0.4484 0.4495 0.4505 0.4515 0.4525 0.4535 0.4545 1.7 0.4554 0.4564 0.4573 0.4582 0.4591 0.4599 0.4608 0.4616 0.4625 0.4633 1.8 0.4641 0.4649 0.4656 0.4664 0.4671 0.4678 0.4686 0.4693 0.4699 0.4706 1.9 0.4713 0.4719 0.4726 0.4732 0.4738 0.4744 0.4750 0.4756 0.4761 0.47672.0 0.4472 0.4778 0.4783 0.4788 0.04793 0.4798 0.4803 0.4808 0.4812 0.4817 2.1 0.4821 0.4826 0.4830 0.4834 0.4838 0.4842 0.4846 0.4850 0.4854 0.4857 2.2 0.4861 0.4864 0.4868 0.4871 0.4875 0.4878 0.4881 0.4884 0.4887 0.48902.3 0.4893 0.4896 0.4898 0.4901 0.4904 0.4906 0.4909 0.4911 0.4913 0.49162.4 0.4918 0.4920 0.4922 0.4925 0.4927 0.4929 0.4931 0.4932 0.4934 0.49362.5 0.4938 0.4940 0.4941 0.4943 0.4945 0.4946 0.4948 0.4949 0.4951 0.4952 2.6 0.4953 0.4955 0.4956 0.4957 0.4959 0.4960 0.4961 0.4962 0.4963 0.4964 2.7 0.4965 0.4966 0.4967 0.4968 0.4969 0.4970 0.4971 0.4972 0.4973 0.49742.8 0.4974 0.4975 0.4976 0.4977 0.4977 0.4978 0.4979 0.4979 0.4980 0.49812.9 0.4981 0.4982 0.4982 0.4983 0.4984 0.4984 0.4985 0.4986 0.4986 0.49863.0 0.4987 0.4987 0.4987 0.4988 0.4988 0.4989 0.4989 0.4989 0.4990 0.4990 x 0.0 0.2 0.4 0.6 0.8 1.0 0.3413447 0.3849303 0.4192433 0.4452007 0.4640697 2.0 0.4772499 0.4860966 0.4918025 0.4953388 0.4974449 3.0 0.4986501 0.4993129 0.4998409 0.4999277 0.49992774.0 0.4999683 0.4999867 0.4999946 0.4999979 0.4999992 /C42 This table is reproduced, by permission, from the Biometrica Tables for Statisticians , vol. 1, 1954, edited by E. S. Pearson and H. O. Hartley and published by the Cambridge University Press for the Biometrica Trustees. /C70urther reading Anton, Howard, Elementary Linear Algebra , 3rd ed., John Wiley, New York, 1982. Arfken, G. B., Weber, H. J., Mathematical Methods for Physicists , 4th ed., Academic Press, New York, 1995. Boas, Mary L., Mathematical Methods in the Physical Sciences , 2nd ed., John Wiley, New York, 1983. Butkov, Eugene, Mathematical Physics , Addison-Wesley, Reading (MA), 1968. Byon, F. W., Fuller, R. W., Mathematics of /C67lassical and /C81uantum Physics , Addison- Wesley, Reading (MA), 1968. Churchill, R. V., Brown, J. W., Verhey, R. F., /C67omplex /C86ariables /C38 Applications , 3rd ed., McGraw-Hill, New York, 1976. Harper, Charles, Introduction to Mathematical Physics , Prentice Hall, Englewood Cli/C128s, NJ, 1976. Kreyszig, E., Advanced Engineering Mathematics , 3rd ed., John Wiley, New York, 1972. Joshi, A. W., Matrices and Tensor in Physics , John Wiley, New York, 1975. Joshi, A. W., Elements of /C71roup Theory for Physicists , John Wiley, New York, 1982. Lass, Harry, /C86ector and Tensor Analysis , McGraw-Hill, New York, 1950. Margenus, Henry, Murphy, George M., The Mathematics of Physics and /C67hemistry ,D . Van Nostrand, New York, 1956. Mathews, Fon, Walker, R. L., Mathematical Methods of Physics , W. A. Benjamin, New York, 1965. Spiegel, M. R., Advanced Mathematics for Engineers and Scientists , Schaum’s Outline Series, McGraw-Hill, New York, 1971. Spiegel, M. R., Theory and Problems of /C86ector Analysis , Schaum’s Outline Series, McGraw-Hill, New York, 1959. Wallace, P. R., Mathematical Analysis of Physical Problems , Dover, New York, 1984. Wong, Chun Wa, Introduction to Mathematical Physics/C44 Methods and /C67oncepts , Oxford, New York, 1991. Wylie, C., Advanced Engineering Mathematics , 2nd ed., McGraw-Hill, New York, 1960. 549 Index Abel’s integral equation, 426 Abelian group, 431 adjoint operator, 212 analytic functions, 243 Cauchy integral formula and, 244Cauchy integral theorem and, 257 angular momentum operator, 18Argand diagram, 234associated Laguerre equation, polynomials see Laguerre equation; Laguerre functions associated Legendre equation, functions see Legendre equation; Legendre functions associated tensors, 53auxiliary (characteristic) equation, 75axial vector, 8 Bernoulli equation, 72 Bessel equation, 321 series solution, 322 Bessel functions, 323 approximations, 335 first kind /C74 n…x†, 324 generating function, 330 hanging flexible chain, 328 Hankel functions, 328integral representation, 331orthogonality, 336recurrence formulas, 332second kind Y n…x†seeNeumann functions spherical, 338 beta function, 95branch line (branch cut), 241branch points, 241 calculus of variations, 347–371 brachistochrone problem, 350 canonical equations of motion, 361constraints, 353 Euler–Lagrange equation, 348 Hamilton’s principle, 361calculus of variations ( contd ) Hamilton–Jacobi equation, 364 Lagrangian equations of motion, 355 Lagrangian multipliers, 353modified Hamilton’s principle, 364 Rayleigh–Ritz method, 359 cartesian coordinates, 3 Cauchy principal value, 289 Cauchy–Riemann conditions, 244 Cauchy’s integral formula, 260 Cauchy’s integral theorem, 257 Cayley–Hamilton theorem, 134change of basis, 224coordinate system, 11 interval, 152 characteristic equation, 125 Christo/C128el symbol, 54 commutator, 107 complex numbers, 233 basic operations, 234polar form, 234roots, 237 connected, simply or multiply, 257contour integrals, 255 contraction of tensors, 50 contravariant tensor, 49convolution theorem Fourier transforms, 188 coordinate system, see specific coordinate system coset, 439covariant di/C128erentiation, 55 covariant tensor, 49 cross product of vectors seevector product of vectors crossing conditions, 441 curl cartesian, 24curvilinear, 32 cylindrical, 34spherical polar, 35 551 curvilinear coordinates, 27 damped oscillations, 80 De Moivre’s formula, 237 del, 22 formulas involving del, 27 delta function, Dirac, 183 Fourier integral, 183Green’s function and, 192point source, 193 determinants, 538–547di/C128erential equations, 62 first order, 63 exact 67integrating factors, 69separable variables 63 homogeneous, 63numerical solutions, 469second order, constant coecients, 72 complementary functions, 74Frobenius and Fuchs theorem, 86particular integrals, 77 singular points, 86 solution in power series, 85 direct product matrices, 139tensors, 50 direction angles and direction cosines, 3divergence cartesian, 22curvilinear, 30cylindrical, 33spherical polar, 35 dot product of vectors, 5dual vectors and dual spaces, 211 eigenvalues and eigenfunctions of Sturm– Liouville equations, 340 hermitian matrices, 124 orthogonality, 129 real, 128 an operator, 217 entire functions, 247Euler’s linear equation, 83 Fourier series, 144 convergence and Dirichlet conditions, 150 di/C128erentiation, 157Euler–Fourier formulas, 145exponential form of Fourier series, 156Gibbs phenomenon, 150half-range Fourier series, 151integration, 157 interval, change of, 152 orthogonality, 162Parseval’s identity, 153vibrtating strings, 157Fourier transform, 164 convolution theorem, 188 delta function derivation, 183 Fourier integral, 164Fourier series ( contd ) Fourier sine and cosine transforms, 172 Green’s function method, 192head conduction, 179Heisenberg’s uncertainty principle, 173Parseval’s identity, 186solution of integral equation, 421transform of derivatives, 190wave packets and group velocity, 174 Fredholm integral equation, 413 seealso integral equations Frobenius’ method seeseries solution of di/C128erential equations Frobenius–Fuch’s theorem, 86function analytic,entire,harmonic, 247 function spaces, 226 gamma function, 94 gauge transformation, 411Gauss’ law, 391Gauss’ theorem, 37generating function, for associated Laguerre polynomials, 320 Bessel functions, 330 Hermite polynomials, 314Laguerre polynomials, 317Legendre polynomials, 301 Gibbs phenomenon, 150gradient cartesian, 20curvilinear, 29cylindrical, 33spherical polar, 35 Gram–Schmidt orthogonalization, 209Green’s functions, 192, construction of one dimension, 192three dimensions, 405 delta function, 193 Green’s theorem, 43 in the plane, 44 group theory, 430 conjugate clsses, 440cosets, 439cyclic group, 433 rotation matrix, 234, 252special unitary group, SU(2), 232 definitions, 430dihedral groups, 446 generator, 451 homomorphism, 436irreducible representations, 442isomorphism, 435multiplication table, 434Lorentz group, 454 permutation group, 438 orthogonal group SO(3) 552INDE/C88 group theory ( contd ) symmetry group, 446 unitary group, 452 unitary unimodular group SU…n† Hamilton–Jacobi equation, 364Hamilton’s principle and Lagrange equations of motion, 355 Hankel functions, 328 Hankel transforms, 385 harmonic functions, 247Helmholtz theorem, 44Hermite equation, 311Hermite polynomials, 312 generating function, 314 orthogonality, 314 recurrence relations, 313 hermitian matrices, 114 orthogonal eigenvectors, 129real eigenvalues, 128 hermitian operator, 220 completeness of eigenfunctions, 221eigenfunctions, orthogonal, 220eigenvalues, real, 220 Hilbert space, 230Hilbert-Schmidt method of solution, 421homogeneous seelinear equations homomorphism, 436 indicial equation, 87 inertia, moment of, 135 infinity seesingularity, pole essential singularity integral equations, 413 Abel’s equation, 426classical harmonic oscillator, 427di/C128erential equation–integral equation transformation, 419 Fourier transform solution, 421Fredholm equation, 413Laplace transform solution, 420Neumann series, 416quantum harmonic oscillator, 427Schmidt–Hilbert method, 421 separable kernel, 414 Volterra equation, 414 integral transforms, 384 see also Fourier transform, Hankel transform, Laplacetransform, Mellin transform Fourier, 164 Hankel, 385 Laplace, 372Mellin, 385 integration, vector, 35 line integrals, 36surface integrals, 36 interpolation, 1461 inverse operator, uniqueness of, 218 irreducible group representations, 442 Jacobian, 29kernels of integral equations, 414 separable, 414 Kronecker delta, 6 mixed second-rank tensor, 53 Lagrangian, 355 Lagrangian multipliers, 354Laguerre equation, 316 associated Laguerre equation, 320 Laguerre functions, 317 associated Laguerre polynomials, 320generating function, 317orthogonality, 319Rodrigues’ representation, 318 Laplace equation, 389 solutions of, 392 Laplace transform, 372 existence of, 373integration of tranforms, 383inverse transformation, 373solution of integral equation, 420 the first shifting theorem, 378 the second shifting theorem, 379transform of derivatives, 382 Laplacian cartesian, 24cylindrical, 34scalar, 24 spherical polar, 35 Laurent expansion, 274 Legendre equation, 296 associated Legendre equation, 307series solution of Legendre equation, 296 Legendre functions, 299 associated Legendre functions, 308generating function, 301orthogonality, 304recurrence relations, 302Rodrigues’ formula, 299 linear combination, 204linear independence, 204 linear operator, 212 Lorentz gauge condition, 411Lorentz group, 454Lorentz transformation, 455 mapping, 239 matrices, 100 anti-hermitian, 114commutatordefinition, 100diagonalization, 129 direct product, 139 eigenvalues and eigenvectors, 124Hermitian, 114nverse, 111matrix multiplication, 103moment of inertia, 135 orthogonal, 115 and unitary transformations, 121 553INDE/C88 matrices ( contd ) Pauli spin, 142 representation, 226 rotational, 117similarity transformation, 122symmetric and skew-symmetric, 109trace, 121transpose, 108unitary, 116 Maxwell equations, 411 derivation of wave equation, 411 Mellin transforms, 385metric tensor, 51mixed tensor, 49Morera’s theorem, 259 multipoles, 248 Neumann functions, 327 Newton’s root finding formula, 465normal modes of vibrations 136numerical methods, 459 roots of equations, 460 false position (linear interpolation), 461graphic methods, 460 Newton’s method, 464 integration, 466 rectangular rule, 455 Simpson’s rule, 469trapezoidal rule, 467 interpolation, 459 least-square fit, 477 solutions of di/C128erential equations, 469 Euler’s rule, 470Runge–Kutta method, 473Taylor series method, 472 system of equations, 476 operators adjoint, 212 angular momentum operator, 18 commuting, 225del, 22di/C128erential operator /C68(ˆd=dx), 78 linear, 212 orthonormal, 161, 207 oscillator, damped, 80 integral equations for, 427simple harmonic, 427 Parseval’s identity, 153, 186partial di/C128erential equations, 387 linear second order, 388 elliptic, hyperbolic, parabolic, 388Green functions, 404Laplace transformationseparation of variables Laplace’s equation, 392, 395, 398 wave equation, 402 Pauli spin matrices, 142phase of a complex number, 235 Poisson equation. 389 polar form of complex numbers, 234 poles, 248probability theory , 481 combinations, 485continuous distributions, 500 Gaussian, 502Maxwell–Boltzmann, 503 definition of probability, 481expectation and variance, 490fundamental theorems, 486probability distributions, 491 binomial, 491Gaussian, 497 Poisson, 495 sample space, 482 power series, 269 solution of di/C128erential equations, 85 projection operators, 222 pseudovectors, 8 quotient rule, 50rank (order), of tensor, 49 of group, 430 Rayleigh–Ritz (variational) method, 359 recurrence relations Bessel functions, 332Hermite functions, 313Laguerre functions, 318Legendre functions, 302 residue theorem, 282 residues 279 calculus of residues, 280 Riccati equation, 98Riemann surface, 241Rodrigues’ formula Hermite polynomials, 313Laguerre polynomials, 318 associated Laguerre polynomials, 320 Legendre polynomials, 299 associated Legendre polynomials, 308 root diagram, 238rotation groups SO…2†;SO…3†, 450 of coordinates, 11, 117 of vectors, 11–13 Runge–Kutta solution, 473 scalar, definition of, 1 scalar potential, 20, 390, 411 scalar product of vectors, 5Schmidt orthogonalization seeGram–Schmidt orthogonalization Schro /C200dinger wave equation, 427 variational approach, 368 Schwarz–Cauchy inequality, 210secular (characteristic) equation, 125 554INDE/C88 series solution of di/C128erential equations, Bessel’s equation, 322 Hermite’s equation, 311 Laguerre’s equation, 316Legendre’s equation, 296 associated Legendre’s equation, 307 similarity transformation, 122singularity, 86, 248 branch point, 240 di/C128erential equation, 86 Laurent series, 274on contour of integration, 290 special unitary group, SU…n†, 452 spherical polar coordinates, 34step function, 380 Stirling’s asymptotic formula for n/C33, 99 Stokes’ theorem, 40 Sturm–Liouville equation, 340subgroup, 439summation convention (Einstein’s), 48symbolic software, 492 Taylor series of elementary functions, 272 tensor analysis, 47 associated tensor, 53basic operations with tensors, 49contravariant vector, 48covariant di/C128erentiation, 55 covariant vector, 48tensor analysis ( contd ) definition of second rank tensor, 49geodesic in Riemannian space, 53 metric tensor, 51quotient law, 50symmetry–antisymmetry, 50 trace (matrix), 121triple scalar product of vectors, 10triple vector product of vectors, 11 uncertainty principle in quantum theory, 173 unit group element, 430unit vectors cartesian, 3cylindrical, 32 spherical polar, 34 variational principles, seecalculus of variations vector and tensor analysis, 1–56 vector potential, 411vector product of vectors, 7 vector space, 13, 199 Volterra integral equation. seeintegral equations wave equation, 389 derivation from Maxwell’s equations, 411solution of, separation of variables, 402 wave packets, 174 group velocity, 174 555INDE/C88