Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / Physics Book Downloads / Math Methods in Physics Books / PDF Originals

Stone M. Methods of Mathematical Physics II (2002)(316s)-1

PDF · 455 pages · 3.1 MB
Open PDF file

Graduate lecture notes by Michael Stone (University of Illinois), from the second half of a two-semester mathematical methods course for first-year physics students. Parts cover tensors, exterior calculus, integration on manifolds, topology, group representations, Lie groups, fibre bundles, complex analysis, and special functions such as the gamma and elliptic functions. This is a downloaded textbook, not Phil's own writing.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Mathematics for Physics II A set of lecture notes by Michael Stone PIMANDER-CASAUBON Alexandria•Florence•London ii Copyright c/circlecopyrt2001,2002,2003 M. Stone. All rights reserved. No part of this material can be reproduc ed, stored or transmitted without the written permission of the author. F or information contact: Michael Stone, Loomis Laboratory of Physics, Univ ersity of Illinois, 1110 West Green Street, Urbana, IL 61801, USA. Preface These notes cover the material from the second half of a two-s emester se- quence of mathematical methods courses given to first year ph ysics graduate students at the University of Illinois. They consist of thre e loosely connected parts: i) an introduction to modern “calculus on manifolds” , the exterior differential calculus, and algebraic topology; ii) an intro duction to group rep- resentation theory and its physical applications; iii) a fa irly standard course on complex variables. iii iv PREFACE Contents Preface iii 1 Tensors in Euclidean Space 1 1.1 Covariant and Contravariant Vectors . . . . . . . . . . . . . . 1 1.2 Tensors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 1.3 Cartesian Tensors . . . . . . . . . . . . . . . . . . . . . . . . . 18 1.4 Further Exercises and Problems . . . . . . . . . . . . . . . . . 29 2 Differential Calculus on Manifolds 33 2.1 Vector and Covector Fields . . . . . . . . . . . . . . . . . . . . 33 2.2 Differentiating Tensors . . . . . . . . . . . . . . . . . . . . . . 39 2.3 Exterior Calculus . . . . . . . . . . . . . . . . . . . . . . . . . 48 2.4 Physical Applications . . . . . . . . . . . . . . . . . . . . . . . 54 2.5 Covariant Derivatives . . . . . . . . . . . . . . . . . . . . . . . 63 2.6 Further Exercises and Problems . . . . . . . . . . . . . . . . . 70 3 Integration on Manifolds 75 3.1 Basic Notions . . . . . . . . . . . . . . . . . . . . . . . . . . . 75 3.2 Integrating p-Forms . . . . . . . . . . . . . . . . . . . . . . . . 79 3.3 Stokes’ Theorem . . . . . . . . . . . . . . . . . . . . . . . . . 84 3.4 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87 3.5 Exercises and Problems . . . . . . . . . . . . . . . . . . . . . . 105 4 An Introduction to Topology 115 4.1 Homeomorphism and Diffeomorphism . . . . . . . . . . . . . . 116 4.2 Cohomology . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 4.3 Homology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122 4.4 De Rham’s Theorem . . . . . . . . . . . . . . . . . . . . . . . 138 v vi CONTENTS 4.5 Poincar´ e Duality . . . . . . . . . . . . . . . . . . . . . . . . . 142 4.6 Characteristic Classes . . . . . . . . . . . . . . . . . . . . . . . 147 4.7 Hodge Theory and the Morse Index . . . . . . . . . . . . . . . 154 5 Groups and Group Representations 171 5.1 Basic Ideas . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171 5.2 Representations . . . . . . . . . . . . . . . . . . . . . . . . . . 179 5.3 Physics Applications . . . . . . . . . . . . . . . . . . . . . . . 192 5.4 Further Exercises and Problems . . . . . . . . . . . . . . . . . 201 6 Lie Groups 207 6.1 Matrix Groups . . . . . . . . . . . . . . . . . . . . . . . . . . 207 6.2 Geometry of SU(2) . . . . . . . . . . . . . . . . . . . . . . . . 213 6.3 Lie Algebras . . . . . . . . . . . . . . . . . . . . . . . . . . . . 234 6.4 Further Exercises and Problems . . . . . . . . . . . . . . . . . 253 7 The Geometry of Fibre Bundles 257 7.1 Fibre Bundles . . . . . . . . . . . . . . . . . . . . . . . . . . 257 7.2 Physics Examples . . . . . . . . . . . . . . . . . . . . . . . . . 259 7.3 Working in the Total Space . . . . . . . . . . . . . . . . . . . 274 8 Complex Analysis I 291 8.1 Cauchy-Riemann equations . . . . . . . . . . . . . . . . . . . . 291 8.2 Complex Integration: Cauchy and Stokes . . . . . . . . . . . . 30 3 8.3 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . 312 8.4 Applications of Cauchy’s Theorem . . . . . . . . . . . . . . . . 318 8.5 Meromorphic functions and the Winding-Number . . . . . . . 3 34 8.6 Analytic Functions and Topology . . . . . . . . . . . . . . . . 337 8.7 Further Exercises and Problems . . . . . . . . . . . . . . . . . 353 9 Complex Analysis II 359 9.1 Contour Integration Technology . . . . . . . . . . . . . . . . . 359 9.2 The Schwarz Reflection Principle . . . . . . . . . . . . . . . . 370 9.3 Partial-Fraction and Product Expansions . . . . . . . . . . . . 381 9.4 Wiener-Hopf Equations II . . . . . . . . . . . . . . . . . . . . 387 9.5 Further Exercises and Problems . . . . . . . . . . . . . . . . . 396 CONTENTS vii 10 Special Functions II 401 10.1 The Gamma Function . . . . . . . . . . . . . . . . . . . . . . 401 10.2 Linear Differential Equations . . . . . . . . . . . . . . . . . . . 40 6 10.3 Solving ODE’s via Contour integrals . . . . . . . . . . . . . . 41 4 10.4 Asymptotic Expansions . . . . . . . . . . . . . . . . . . . . . . 421 10.5 Elliptic Functions . . . . . . . . . . . . . . . . . . . . . . . . . 432 10.6 Further Exercises and Problems . . . . . . . . . . . . . . . . . 439 viii CONTENTS Chapter 1 Tensors in Euclidean Space In this chapter we explain how a vector space Vgives rise to a family of associated tensor spaces, and how mathematical objects suc h as linear maps or quadratic forms should be understood as being elements of these spaces. We then apply these ideas to physics. We make extensive use of notions and notations from the appendix on linear algebra, so it may help to review that material before we begin. 1.1 Covariant and Contravariant Vectors When we have a vector space VoverR, and{e1,e2,...,en}and{e/prime 1,e/prime 2,...,e/prime n} are both bases for V, then we may expand each of the basis vectors eµin terms of the e/prime µas eν=aµ νe/prime µ. (1.1) We are here, as usual, using the Einstein summation conventi on that repeated indices are to be summed over. Written out in full for a three- dimensional space, the expansion would be e1=a1 1e/prime 1+a2 1e/prime 2+a3 1e/prime 3, e2=a1 2e/prime 1+a2 2e/prime 2+a3 2e/prime 3, e3=a1 3e/prime 1+a2 3e/prime 2+a3 3e/prime 3. We could also have expanded the e/prime µin terms of the eµas e/prime ν= (a−1)µ νe/prime µ. (1.2) 1 2 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE As the notation implies, the matrices of coefficients aµ νand (a−1)µ νare inverses of each other: aµ ν(a−1)ν σ= (a−1)µ νaν σ=δµ σ. (1.3) If we know the components xµof a vector xin theeµbasis then the compo- nentsx/primeµofxin the e/prime µbasis are obtained from x=x/primeµe/prime µ=xνeν= (xνaµ ν)e/prime µ (1.4) by comparing the coefficients of e/prime µ. We find that x/primeµ=aµ νxν. Observe how theeµand thexµtransform in “opposite” directions. The components xµ are therefore said to transform contra variantly . Associated with the vector space Vis itsdual spaceV∗, whose elements arecovectors ,i.e.linear maps f:V→R. Iff∈V∗andx=xµeµ, we use the linearity property to evaluate f(x) as f(x) =f(xµeµ) =xµf(eµ) =xµfµ. (1.5) Here, the set of numbers fµ=f(eµ) are the components of the covector f. If we change basis so that eν=aµ νe/prime µthen fν=f(eν) =f(aµ νe/prime µ) =aµ νf(e/prime µ) =aµ νf/prime µ. (1.6) We conclude that fν=aµ νf/prime µ. Thefµcomponents transform in the same man- ner as the basis. They are therefore said to transform covariantly . In physics it is traditional to call the the set of numbers xµwith upstairs indices (the components of) a contravariant vector . Similarly, the set of numbers fµwith downstairs indices is called (the components of) a covariant vector . Thus, contravariant vectors are elements of Vand covariant vectors are elements ofV∗. The relationship between VandV∗is one of mutual duality, and to mathematicians it is only a matter of convenience which spac e isVand which space is V∗. The evaluation of f∈V∗onx∈Vis therefore often written as a “pairing” ( f,x), which gives equal status to the objects being put togther to get a number. A physics example of such a mutual ly dual pair is provided by the space of displacements xand the space of wave-numbers k. The units of xandkare different (meters versus meters−1). There is therefore no meaning to “ x+k,” and xandkare not elements of the same vector space. The “dot” in expressions such as ψ(x) =eik·x(1.7) 1.1. COVARIANT AND CONTRAVARIANT VECTORS 3 cannot be a true inner product (which requires the objects it links to be in the same vector space) but is instead a pairing (k,x)≡k(x) =kµxµ. (1.8) In describing the physical world we usually give priority to the space in which we live, breathe and move, and so treat it as being “ V”. The displacement vector xthen becomes the contravariant vector, and the Fourier-spa ce wave- number k, being the more abstract quantity, becomes the covariant co vector. Our vector space may come equipped with a metric that is derived from a non-degenerate inner product. We regard the inner product as being a bilinear form g:V×V→R, so the length/bardblx/bardblof a vector xis/radicalbig g(x,x). The set of numbers gµν=g(eµ,eν) (1.9) comprises the (components of) the metric tensor . In terms of them, the inner of product/angbracketleftx,y/angbracketrightof pair of vectors x=xµeµandy=yµeµbecomes /angbracketleftx,y/angbracketright≡g(x,y) =gµνxµyν. (1.10) Real-valued inner products are always symmetric, so g(x,y) =g(y,x) and gµν=gνµ. As the product is non-degenerate, the matrix gµνhas an inverse, which is traditionally written as gµν. Thus gµνgνλ=gλνgνµ=δλ µ. (1.11) The additional structure provided by the metric permits us t o identifyV withV∗. The identification is possible, because, given any f∈V∗, we can find a vector/tildewidef∈Vsuch that f(x) =/angbracketleft/tildewidef,x/angbracketright. (1.12) We obtain/tildewidefby solving the equation fµ=gµν/tildewidefν(1.13) to get/tildewidefν=gνµfµ. We may now drop the tilde and identify fwith/tildewidef, and henceVwithV∗. When we do this, we say that the covariant components fµare related to the contravariant components fµbyraising fµ=gµνfν, (1.14) 4 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE orlowering fµ=gµνfν, (1.15) the indexµusing the metric tensor. Bear in mind that this V∼=V∗identi- fication depends crucially on the metric. A different metric w ill, in general, identify an f∈V∗with a completely different /tildewidef∈V. We may play this game in the Euclidean space Enwith its “dot” inner product. Given a vector xand a basis eµfor whichgµν=eµ·eν, we can define two sets of components for the same vector. Firstly the coefficients xµ appearing in the basis expansion x=xµeµ, (1.16) and secondly the “components” xµ=eµ·x=g(eµ,x) =g(eµ,xνeν) =g(eµ,eν)xν=gµνxν(1.17) ofxalong the basis vectors. These two set of numbers are then res pectively called the contravariant and covariant components of the ve ctorx. If the eµconstitute an orthonormal basis, where gµν=δµν, then the two sets of components (covariant and contravariant) are numerically coincident. In a non-orthogonal basis they will be different, and we must take care never to add contravariant components to covariant ones. 1.2 Tensors We now introduce tensors in two ways: firstly as sets of number s labelled by indices and equipped with transformation laws that tell us h ow these numbers change as we change basis; and secondly as basis-independen t objects that are elements of a vector space constructed by taking multipl e tensor products of the spaces VandV∗. 1.2.1 Transformation rules After we change basis eµ→e/prime µ, where eν=aµ νe/prime µ, the metric tensor will be represented by a new set of components g/prime µν=g(e/prime µ,e/prime ν). (1.18) 1.2. TENSORS 5 These are be related to the old components by gµν=g(eµ,eν) =g(aρ µe/prime ρ,aσ νe/prime σ) =aρ µaσ νg(e/prime ρ,e/prime σ) =aρ µaσ νg/prime ρσ. (1.19) This transformation rule for gµνhas both of its subscripts behaving like the downstairs indices of a covector. We therefore say that gµνtransforms as a doubly covariant tensor . Written out in full, for a two-dimensional space, the transformation law is g11=a1 1a1 1g/prime 11+a1 1a2 1g/prime 12+a2 1a1 1g/prime 21+a2 1a2 1g/prime 22, g12=a1 1a1 2g/prime 11+a1 1a2 2g/prime 12+a2 1a1 2g/prime 21+a2 1a2 2g/prime 22, g21=a1 2a1 1g/prime 11+a1 2a2 1g/prime 12+a2 2a1 1g/prime 21+a2 2a2 1g/prime 22, g22=a1 2a1 2g/prime 11+a1 2a2 2g/prime 12+a2 2a1 2g/prime 21+a2 2a2 2g/prime 22. In three dimensions each row would have nine terms, and sixte en in four dimensions. We see why Einstein was driven to invent his summ ation con- vention! A set of numbers Qαβ γδ/epsilon1, whose indices range from 1 to the dimension of the space and that transforms as Qαβ γδ/epsilon1= (a−1)α α/prime(a−1)β β/primeaγ/prime γaδ/prime δa/epsilon1/prime /epsilon1Q/primeα/primeβ/prime γ/primeδ/prime/epsilon1/prime, (1.20) or conversely as Q/primeαβ γδ/epsilon1=aα α/primeaβ β/prime(a−1)γ/prime γ(a−1)δ/prime δ(a−1)/epsilon1/prime /epsilon1Qα/primeβ/prime γ/primeδ/prime/epsilon1/prime, (1.21) comprises the components of a doubly contravariant, triply covariant tensor. More compactly, the Qαβ γδ/epsilon1are the components of a tensor of type (2 ,3). Tensors of type ( p,q) are defined analogously. The total number of indices p+qis called the rankof the tensor. Note how the indices are wired up in the transformation rules (1.20) and (1.21): free (not summed over) upstairs indices on the left h and side of the equations match to free upstairs indices on the right hand si de, similarly for the downstairs indices. Also upstairs indices are summed on ly with down- stairs ones. Similar conditions apply to equations relating tensors in a ny particular basis. If they are violated you do not have a valid tensor equa tion — meaning that an equation valid in one basis will not be valid in anothe r basis. Thus an equation Aµ νλ=Bµτ νλτ+Cµ νλ (1.22) 6 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE is fine, but Aµ νλ?=Bν µλ+Cµ νλσσ+Dµ νλτ (1.23) has something wrong in each term. Incidentally, although not illegal, it is a good idea not to w rite tensor indices directly underneath one another — i.e.do not write Qij kjl— because if you raise or lower indices using the metric tensor, and som e pages later in a calculation try to put them back where they were, they might end up in the wrong order. Tensor algebra The sum of two tensors of a given type is also a tensor of that ty pe. The sum of two tensors of different types is not a tensor. Thus each par ticular type of tensor constitutes a distinct vector space, but one derived from the common underlying vector space whose change-of-basis formula is b eing utilized. Tensors can be combined by multiplication: if Aµ νλandBµ νλτare tensors of type (1,2) and (1,3) respectively, then Cαβ νλρστ=Aα νλBβ ρστ (1.24) is a tensor of type (2 ,5). An important operation is contraction , which consists of setting one or more contravariant index index equal to a covariant index an d summing over the repeated indices. This reduces the rank of the tensor. So , for example, Dρστ=Cαβ αβρστ (1.25) is a tensor of type (0 ,3). Similarly f(x) =fµxµis a type (0,0) tensor, i.e.an invariant — a number that takes the same value in all bases. Upper indice s can only be contracted with lower indices, and vice versa . For example, the array of numbers Aα=Bαββobtained from the type (0 ,3) tensorBαβγisnot a tensor of type (0 ,1). The contraction procedure outputs a tensor because setting an upper index and a lower index to a common value µand summing over µ, leads to the factor...(a−1)µ αaβ µ...appearing in the transformation rule. Now (a−1)µ αaβ µ=δβ α, (1.26) and the Kronecker delta effects a summation over the correspo nding pair of indices in the transformed tensor. 1.2. TENSORS 7 Although often associated with general relativity, tensor s occur in many places in physics. They are used, for example, in elasticity theory, where the word “tensor” in its modern meaning was introduced by Woldem ar Voigt in 1898. Voigt, following Cauchy and Green, described the in finitesimal deformation of an elastic body by the strain tensor eαβ, which is a tensor of type (0,2). The forces to which the strain gives rise are de scribed by the stress tensor σλµ. A generalization of Hooke’s law relates stress to strain vi a a tensor of elastic constants cαβγδas σαβ=cαβγδeγδ. (1.27) We study stress and strain in more detail later in this chapte r. Exercise 1.1 : Show that gµν, the matrix inverse of the metric tensor gµν, is indeed a doubly contravariant tensor, as the position of its indices suggests. 1.2.2 Tensor character of linear maps and quadratic forms As an illustration of the tensor concept and of the need to dis tinguish be- tween upstairs and downstairs indices, we contrast the prop erties of matrices representing linear maps and those representing quadratic forms. A linear map M:V→Vis an object that exists independently of any basis. Given a basis, however, it is represented by a matrix Mµνobtained by examining the action of the map on the basis elements: M(eµ) =eνMν µ. (1.28) Acting on xwe get a new vector y=M(x), where yνeν=y=M(x) =M(xµeµ) =xµM(eµ) =xµMν µeν=Mν µxµeν.(1.29) We therefore have yν=Mν µxµ, (1.30) which is the usual matrix multiplication y=Mx. When we change basis, eν=aµ νe/prime µ, then eνMν µ=M(eµ) =M(aρ µe/prime ρ) =aρ µM(e/prime ρ) =aρ µe/prime σM/primeσ ρ=aρ µ(a−1)ν σeνM/primeσ ρ. (1.31) 8 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE Comparing coefficients of eν, we find Mν µ=aρ µ(a−1)ν σM/primeσ ρ, (1.32) or, conversely, M/primeν µ= (a−1)ρ µaν σMσ ρ. (1.33) Thus a matrix representing a linear map has the tensor charac ter suggested by the position of its indices, i.e.it transforms as a type (1 ,1) tensor. We can derive the same formula in matrix notation. In the new basis t he vectors x andyhave new components x/prime=Ax, andy/prime=Ay. Consequently y=Mx becomes y/prime=Ay=AMx =AMA−1x/prime, (1.34) and the matrix representing the map Mhas new components M/prime=AMA−1. (1.35) Now consider the quadratic form Q:V→Rthat is obtained from a symmetric bilinear form Q:V×V→Rby settingQ(x) =Q(x,x). We can write Q(x) =Qµνxµxν=xµQµνxν=xTQx, (1.36) whereQµν≡Q(eµ,eν) are the entries in the symmetric matrix Q, the suffixT denotes transposition, and xTQxis standard matrix-multiplication notation. Just as does the metric tensor, the coefficients Qµνtransform as a type (0 ,2) tensor: Qµν=aα µaβ νQ/prime αβ. (1.37) In matrix notation the vector xagain transforms to have new components x/prime=Ax, butx/primeT=xTAT. Consequently x/primeTQ/primex/prime=xTATQ/primeAx. (1.38) Thus Q=ATQ/primeA. (1.39) The message is that linear maps and quadratic forms can both b e represented by matrices, but these matrices correspond to distinct type s of tensor and transform differently under a change of basis. A matrix representing a linear map has a basis-independent d eterminant. Similarly the traceof a matrix representing a linear map trMdef=Mµ µ (1.40) 1.2. TENSORS 9 is a tensor of type (0 ,0), i.e. a scalar, and therefore basis independent. On the other hand, while you can certainly compute the determin ant or the trace of the matrix representing a quadratic form in some particul ar basis, when you change basis and calculate the determinant or trace of th e transformed matrix, you will get a different number. Itispossible to make a quadratic form out of a linear map, but this requires using the metric to lower the contravariant index o n the matrix representing the map: Q(x) =xµgµνQν λxλ=x·Qx. (1.41) Be careful, therefore: the matrices “ Q” inxTQxand in x·Qxare representing different mathematical objects. Exercise 1.2 : In this problem we will use the distinction between the tran s- formation law of a quadratic form and that of a linear map to re solve the following “paradox”: •In quantum mechanics we are taught that the matrices represe nting two operators can be simultaneously diagonalized only if they c ommute. •In classical mechanics we are taught how, given the Lagrangi an L=/summationdisplay ij/parenleftbigg1 2˙qiMij˙qj−1 2qiVijqj/parenrightbigg , to construct normal co-ordinates Qisuch thatLbecomes L=/summationdisplay i/parenleftbigg1 2˙Q2 i−1 2ω2 iQ2 i/parenrightbigg . We have apparantly managed to simultaneously diagonize the matricesMij→ diag (1,...,1) andVij→diag (ω2 1,...,ω2 n), even though there is no reason for them to commute with each other! Show that when MandVare a pair of symmetric matrices, with Mbeing positive definite, then there exits an invertible matrix Asuch that ATMAand ATVAare simultaneously diagonal. (Hint: Consider Mas defining an inner product, and use the Gramm-Schmidt procedure to first find a or thonormal frame in which M/prime ij=δij. Then show that the matrix corresponding to V in this frame can be diagonalized by a further transformatio n that does not perturb the already diagonal M/prime ij.) 10 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE 1.2.3 Tensor product spaces We may regard the set of numbers Qαβ γδ/epsilon1as being the components of an object Qthat is element of the vector space of type (2 ,3) tensors. We denote this vector space by the symbol V⊗V⊗V∗⊗V∗⊗V∗, the notation indicating that it is derived from the original Vand its dual V∗by taking tensor products of these spaces. The tensor Qis to be thought of as existing as an element of V⊗V⊗V∗⊗V∗⊗V∗independently of any basis, but given a basis{eµ}forV, and the dual basis {e∗ν}forV∗, we expand it as Q=Qαβ γδ/epsilon1eα⊗eβ⊗e∗γ⊗e∗δ⊗e∗/epsilon1. (1.42) Here the tensor product symbol “ ⊗” is distributive a⊗(b+c) =a⊗b+a⊗c, (a+b)⊗c=a⊗c+b⊗c, (1.43) and associative (a⊗b)⊗c=a⊗(b⊗c), (1.44) but is not commutative a⊗b/negationslash=b⊗a. (1.45) Everything commutes with the field, however, λ(a⊗b) = (λa)⊗b=a⊗(λb). (1.46) If we change basis eα=aβ αe/prime βthen these rules lead, for example, to eα⊗eβ=aλ αaµ βe/prime λ⊗e/prime µ. (1.47) From this change-of-basis formula, we deduce that Tαβeα⊗eβ=Tαβaλ αaµ βe/prime λ⊗e/prime µ=T/primeλµe/prime λ⊗e/prime µ, (1.48) where T/primeλµ=Tαβaλ αaµ β. (1.49) The analogous formula for eα⊗eβ⊗e∗γ⊗e∗δ⊗e∗/epsilon1reproduces the transfor- mation rule for the components of Q. The meaning of the tensor product of a collection of vector sp aces should now be clear: If eµconsititute a basis for V, the space V⊗Vis, for example, 1.2. TENSORS 11 the space of all linear combinations1of the abstract symbols eµ⊗eν, which we declare by fiatto constitute a basis for this space. There is no geometric significance (as there is with a vector product a×b) to the tensor product a⊗b, so the eµ⊗eνare simply useful place-keepers. Remember that these areordered pairs,eµ⊗eν/negationslash=eν⊗eµ. Although there is no geometric meaning, it is possible, however, to give analgebraic meaning to a product like e∗λ⊗e∗µ⊗e∗νby viewing it as a multilinear form V×V×V:→R. We define e∗λ⊗e∗µ⊗e∗ν(eα,eβ,eγ) =δλ αδµ βδν γ. (1.50) We may also regard it as a linear map V⊗V⊗V:→Rby defining e∗λ⊗e∗µ⊗e∗ν(eα⊗eβ⊗eγ) =δλ αδµ βδν γ (1.51) and extending the definition to general elements of V⊗V⊗Vby linearity. In this way we establish an isomorphism V∗⊗V∗⊗V∗∼=(V⊗V⊗V)∗. (1.52) This multiple personality is typical of tensor spaces. We ha ve already seen that the metric tensor is simultaneously an element of V∗⊗V∗and a map g:V→V∗. Tensor products and quantum mechanics When we have two quantum-mechanical systems having Hilbert spacesH(1) andH(2), the Hilbert space for the combined system is H(1)⊗H(2). Quantum mechanics books usually denote the vectors in these spaces b y the Dirac “bra- ket” notation in which the basis vectors of the separate spac es are denoted by2|n1/angbracketrightand|n2/angbracketright, and that of the combined space by |n1,n2/angbracketright. In this notation, a state in the combined system is a linear combination |Ψ/angbracketright=/summationdisplay n1,n2|n1,n2/angbracketright/angbracketleftn1,n2|Ψ/angbracketright, (1.53) 1Do not confuse the tensor-product space V⊗Wwith the Cartesian product V×W. The latter is the set of all ordered pairs ( x,y),x∈V,y∈W. The tensor product includes alsoformal sums of such pairs. The Cartesian product of two vector spaces can be given the structure of a vector space by defining an addition operat ionλ(x1,y1) +µ(x2,y2) = (λx1+µx2,λy1+µy2), but this construction does not lead to the tensor product. Instead it defines the direct sum V⊕W. 2We assume for notational convenience that the Hilbert space s are finite dimensional. 12 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE This is the tensor product in disguise. To unmask it, we simpl y make the notational translation |Ψ/angbracketright → Ψ /angbracketleftn1,n2|Ψ/angbracketright →ψn1,n2 |n1/angbracketright → e(1) n1 |n2/angbracketright → e(2) n2 |n1,n2/angbracketright → e(1) n1⊗e(2) n2. (1.54) Then (1.53) becomes Ψ=ψn1,n2e(1) n1⊗e(2) n2. (1.55) Entanglement: Suppose thatH(1)has basis e(1) 1,...,e(1) mandH(2)has basis e(2) 1,...,e(2) n. The Hilbert space H(1)⊗H(2)is thennmdimensional. Consider a state Ψ=ψije(1) i⊗e(2) j∈H(1)⊗H(2). (1.56) If we can find vectors Φ≡φie(1) i∈H(1), X≡χje(2) j∈H(2), (1.57) such that Ψ=Φ⊗X≡φiχje(1) i⊗e(2) j (1.58) then the tensor Ψis said to be decomposable and the two quantum systems are said to be unentangled . If there are no such vectors then the two systems areentangled in the sense of the Einstein-Podolski-Rosen (EPR) paradox. Quantum states are really in one-to-one correspondence wit hraysin the Hilbert space, rather than vectors. If we denote the ndimensional vector space over the field of the complex numbers as Cn, the space of rays, in which we do not distinguish between the vectors xandλxwhenλ/negationslash= 0, is denoted byCPn−1and is called complex projective space . Complex projective space is where algebraic geometry is studied. The set of decomposable states may be thought of as a subset of the complex projective space CPnm−1, and, since, as the following excercise shows, this subset is defined by a fi nite number of homogeneous polynomial equations, it forms what algebraic geometers call a variety . This particular subset is known as the Segre variety . 1.2. TENSORS 13 Exercise 1.3 : The Segre conditions for a state to be decomposable: i) By counting the number of independent components that are at our dis- posal in Ψ, and comparing that number with the number of free param- eters in Φ⊗X, show that the coefficients ψijmust satisfy ( n−1)(m−1) relations if the state is to be decomposable. ii) If the state is decomposable, show that 0 =/vextendsingle/vextendsingle/vextendsingle/vextendsingleψijψil ψkjψkl/vextendsingle/vextendsingle/vextendsingle/vextendsingle for all sets of indices i,j,k,l . iii) Assume that ψ11is not zero. Using your count from part (i) as a guide, find a subset of the relations from part (ii) that constitute a necessary and sufficient set of conditions for the state Ψ to be decomposable . Include a proof that your set is indeed sufficient. 1.2.4 Symmetric and skew-symmetric tensors By examining the transformation rule you may see that if a pai r of up- stairs or downstairs indices is symmetric (sayQµν ρστ=Qνµ ρστ) orskew- symmetric (Qµν ρστ=−Qνµ ρστ) in one basis, it remains so after the basis has been changed. (This is nottrue of a pair composed of one upstairs and one downstairs index.) It makes sense, therefore, to defi ne symmetric and skew-symmetric tensor product spaces. Thus skew-symme tric doubly- contravariant tensors can be regarded as belonging to the sp ace denoted by/logicalandtext2Vand expanded as A=1 2Aµνeµ∧eν, (1.59) where the coefficients are skew-symmetric, Aµν=−Aνµ, and the wedge prod- uctof the basis elements is associative and distributive, as is the tensor product, but in addition obeys eµ∧eν=−eν∧eµ. The “1/2” (replaced by 1/p! when there are pindices) is convenient in that each independent component only appears once in the sum. For example, in three dimensions, 1 2Aµνeµ∧eν=A12e1∧e2+A23e2∧e3+A31e3∧e1. (1.60) Symmetric doubly-contravariant tensors can be regarded as belonging to the space sym2Vand expanded as S=Sαβeα⊙eβ (1.61) 14 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE where eα⊙eβ=eβ⊙eαandSαβ=Sβα. (We do not insert a “1/2” here because including it leads to no particular simplification i n any consequent equations.) We can treat these symmetric and skew-symmetric products as symmetric or skew multilinear forms. Define, for example, e∗α∧e∗β(eµ,eν) =δα µδβ ν−δα νδβ µ, (1.62) and e∗α∧e∗β(eµ∧eν) =δα µδβ ν−δα νδβ µ. (1.63) We need two terms on the right-hand-side of these examples be cause the skew-symmetry of e∗α∧e∗β(,) in its slots does not allow us the luxury of demanding that the eµbe inserted in the exact order of the e∗αto get a non-zero answer. Because the p-th order analogue of (1.62) form has p! terms on its right-hand side, some authors like to divide the right -hand-side by p! in this definition. We prefer the one above, though. With our d efinition, and withA=1 2Aµνe∗µ∧e∗νandB=1 2Bαβeα∧eβ, we have A(B) =1 2AµνBµν=/summationdisplay µ<νAµνBµν, (1.64) so the sum is only over independent terms. The wedge (∧) product notation is standard in mathematics wherever skew-symmetry is implied.3The “sym” and⊙are not. Different authors use different notations for spaces of symmetric tensors. This re flects the fact that skew-symmetric tensors are extremely useful and appear in m any different parts of mathematics, while symmetric ones have fewer speci al properties (although they are common in physics). Compare the relative usefulness of determinants and permanents. Exercise 1.4 : Show that in ddimensions: i) the dimension of the space of skew-symmetric covariant te nsors with p indices isd!/p!(d−p)!; ii) the dimension of the space of symmetric covariant tensor s withpindices is (d+p−1)!/p!(d−1)!. 3Skew products, along with the first formulation of the idea of an abstract vector space, were introduced in Hermann Grassmann’s Ausdehnungslehre (1844). Grassmann’s mathematics was not appreciated in his lifetime. In his disa ppointment he turned to other fields, making significant contributions to the theory of col our mixtures (Grassmann’s law), and to the philology of Indo-European languages (anot her Grassmann’s law). 1.2. TENSORS 15 Bosons and fermions Spaces of symmetric and skew-symmetric tensors appear when ever we deal with the quantum mechanics of many indistinguishable parti cles possessing Bose or Fermi statistics. If we have a Hilbert space Hof single-particle states with basis eithen theN-boson space is SymNHwhich consists of states Φ= Φi1i2...iNei1⊙ei2⊙···⊙ eiiN, (1.65) and theN-fermion space is/logicalandtextNH, which contains states Ψ=1 N!Ψi1i2...iNei1∧ei2∧···∧ eiN. (1.66) The symmetry of the Bose wavefunction Φi1...iα...iβ...iN= Φi2...iβ...iα...iN, (1.67) and the skew-symmetry of the Fermion wavefunction Ψi1...iα...iβ...iN=−Ψi2...iβ...iα...iN, (1.68) under the interchange of the particle labels α,βis then natural. Slater Determinants and the Pl¨ ucker Relations : SomeN-fermion states can be decomposed into a product of single-particle states Ψ=ψ1∧ψ2∧···∧ψN =ψi1 1ψi2 2···ψiN Nei1∧ei2∧···∧ eiN. (1.69) Comparing the coefficients of ei1∧ei2∧···∧ eiNin (1.66) and (1.69) shows that the many-body wavefunction can then be written as Ψi1i2...iN=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleψi1 1ψi2 1···ψiN 1 ψi1 2ψi2 2···ψiN 2............ ψi1 Nψi2 N···ψiN N/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle. (1.70) The wavefunction is therefore given by a single Slater determinant . Such wavefunctions correspond to a very special class of states. The general many-fermion state is not decomposable, and its wavefuncti on can only be expressed as a sum of many Slater determinants. The Hartree- Fock method 16 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE of quantum chemistry is a variational approximation that ta kes such a single Slater determinant as its trial wavefunction and varies onl y the one-particle wavefunctions/angbracketlefti|ψa/angbracketright ≡ψi a. It is a remarkably successful approximation, given the very restricted class of wavefunctions it explore s. As with the Segre condition for two distinguishable quantum systems to be unentangled, there is a set of necessary and sufficient cond itions on the Ψi1i2...iNfor the state Ψto be decomposable into single-particle states. The conditions are that Ψi1i2...iN−1[j1Ψj1j2...jN+1]= 0 (1.71) for any choice of indices i1,...iN−1andj1,...,jN+1. The square brackets [...] indicate that the expression is to be antisymmetrized over the indices enclosed in the brackets. For example, a three-particle sta te is decomposable if and only if Ψi1i2j1Ψj2j3j4−Ψi1i2j2Ψj1j3j4+ Ψi1i2j3Ψj1j2j4−Ψi1i2j4Ψj1j2j3= 0.(1.72) These conditions are called the Pl¨ ucker relations after Julius Pl¨ ucker who discovered them long before before the advent of quantum mec hanics.4It is easy to show that Pl¨ ucker’s relations are necessary condit ions for decompos- ability. It takes more sophistication to show that they are s ufficient. We will therefore defer this task to the exercises as the end of the ch apter. As far as we are aware, the Pl¨ ucker relations are not exploited by qua ntum chemists, but, in disguise as the Hirota bilinear equations , they constitute the geometric condition underpinning the many-soliton solutions of the K orteweg-de-Vries and other soliton equations. 1.2.5 Kronecker and Levi-Civita tensors Suppose the tensor δµ νis defined, with respect to some basis, to be unity if µ=νand zero otherwise. In a new basis it will transform to δ/primeµ ν=aµ ρ(a−1)σ νδρ σ=aµ ρ(a−1)ν ρ=δµ ν. (1.73) In other words the Kronecker delta symbol of type (1 ,1) has the same numer- ical components in all co-ordinate systems. This is not true of the Kroneker delta symbol of type (0 ,2),i.e.ofδµν. 4As well as his extensive work in algebraic geometry, Pl¨ ucke r (1801-68) made important discoveries in experimental physics. He was, for example, t he first person to observe the deflection of cathode rays — beams of electrons — by a magnetic field, and the first to point out that each element had its characteristic emission spectrum. 1.2. TENSORS 17 Now consider an n-dimensional space with a tensor ηµ1µ2...µnwhose com- ponents, in some basis, coincides with the Levi-Civita symb ol/epsilon1µ1µ2...µn. We find that in a new frame the components are η/prime µ1µ2...µn= (a−1)ν1 µ1(a−1)ν2 µ2···(a−1)νn µn/epsilon1ν1ν2...νn =/epsilon1µ1µ2...µn(a−1)ν1 1(a−1)ν2 2···(a−1)νn n/epsilon1ν1ν2...νn =/epsilon1µ1µ2...µndetA−1 =ηµ1µ2...µndetA−1. (1.74) Thus, unlike the δµ ν, the Levi-Civita symbol is not quite a tensor. Consider also the quantity √gdef=/radicalBig det [gµν]. (1.75) Here we assume that the metric is positive-definite, so that t he square root is real, and that we have taken the positive square root. Sinc e det [g/prime µν] = det [(a−1)ρ µ(a−1)σ νgρσ] = (det A)−2det [gµν], (1.76) we see that/radicalbig g/prime=|detA|−1√g (1.77) Thus√gis also not quite an invariant. This is only to be expected, be cause g(,) is a quadratic form and we know that there is no basis-indepe ndent meaning to the determinant of such an object. Now define εµ1µ2...µn=√g/epsilon1µ1µ2...µn, (1.78) and assume that εµ1µ2...µnhas the type (0 ,n) tensor character implied by its indices. When we look at how this transforms, and restric t ourselves toorientation preserving changes of of bases, i.e.ones for which det Ais positive, we see that factors of det Aconspire to give ε/prime µ1µ2...µn=/radicalbig g/prime/epsilon1µ1µ2...µn. (1.79) A similar exercise indictes that if we define /epsilon1µ1µ2...into be numerically equal to/epsilon1i1i2...µnthen εµ1µ2...µn=1√g/epsilon1µ1µ2...µn(1.80) 18 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE also transforms as a tensor — in this case a type ( n,0) contravariant one — provided that the factor of 1 /√gis always calculated with respect to the current basis. If the dimension nis even and we are given a skew-symetric tensor Fµν, we can therefore construct an invariant εµ1µ2...µnFµ1µ2···Fµn−1µn=1√g/epsilon1µ1µ2...µnFµ1µ2···Fµn−1µn. (1.81) Similarly, given an skew-symmetric covariant tensor Fµ1...µmwithm(≤n) indices we can form its dual, denoted by F∗, a (n−m)-contravariant tensor with components (F∗)µm−1...µn=1 m!εµ1µ2...µnFµ1...µm=1√g1 m!/epsilon1µ1µ2...µnFµ1...µm. (1.82) We meet this “dual” tensor again, when we study differential f orms. 1.3 Cartesian Tensors If we restrict ourselves to Cartesian co-ordinate systems h aving orthonormal basis vectors, so that gij=δij, then there are considerable simplifications. In particular, we do not have to make a distinction between co - and contra- variant indices. We shall usually write their indices as rom an-alphabet suf- fixes. A change of basis from one orthogonal n-dimensional basis eito another e/prime iwill set e/prime i=Oijej, (1.83) where the numbers Oijare the entries in an orthogonal matrix O,i.e.a real matrix obeying OTO=OOT=I, whereTdenotes the transpose. The set ofn-by-northogonal matrices constitutes the orthogonal group O(n). 1.3.1 Isotropic tensors The Kronecker δijwith both indices downstairs is unchanged by O( n) trans- formations, δ/prime ij=OikOjlδkl=OikOjk=OikOT kj=δij, (1.84) 1.3. CARTESIAN TENSORS 19 and has the same components in any Cartesian frame. We say tha t its components are numerically invariant . A similar property holds for tensors made up of products of δij, such as Tijklmn =δijδklδmn. (1.85) It is possible to show5that any tensor whose components are numerically invariant under all orthogonal transformations is a sum of p roducts of this form. The most general O( n) invariant tensor of rank four is, for example. αδijδkl+βδikδlj+γδilδjk. (1.86) The determinant of an orthogonal transformation must be ±1. If we only allow orientation-preserving changes of basis then we rest rict ourselves to orthogonal transformations Oijwith det O= 1. These are the proper or- thogonal transformations. In ndimensions they constitute the group SO( n). Under SO(n) transformations, both δijand/epsilon1i1i2...inare numerically invariant and the most general SO( n) invariant tensors consist of sums of products of δij’s and/epsilon1i1i2...in’s. The most general SO(4)-invariant rank-four tensor is, f or example, αδijδkl+βδikδlj+γδilδjk+λ/epsilon1ijkl. (1.87) Tensors that are numerically invariant under SO( n) are known as isotropic tensors . As there is no longer any distinction between co- and contrav ariant in- dices, we can now contract any pair of indices. In three dimen sions, for example, Bijkl=/epsilon1nij/epsilon1nkl (1.88) is a rank-four isotropic tensor. Now /epsilon1i1...inisnotinvariant when we transform via an orthogonal transformation with det O=−1, but the product of two /epsilon1’sisinvariant under such transformations. The tensor Bijklis therefore numerically invariant under the larger group O(3) and must b e expressible as Bijkl=αδijδkl+βδikδlj+γδilδjk (1.89) for some coefficients α,βandγ. The following exercise explores some con- sequences of this and related facts. 5The proof is surprisingly complicated. See, for example, M. Spivak, A Comprehensive Introduction to Differential Geometry (second edition) Vol. V, pp. 466-481. 20 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE Exercise 1.5 : We defined the n-dimensional Levi-Civita symbol by requiring that/epsilon1i1i2...inbe antisymmetric in all pairs of indices, and /epsilon112...n= 1. a) Show that /epsilon1123=/epsilon1231=/epsilon1312, but that/epsilon11234=−/epsilon12341=/epsilon13412=−/epsilon14123. b) Show that /epsilon1ijk/epsilon1i/primej/primek/prime=δii/primeδjj/primeδkk/prime+ five other terms, where you should write out all six terms explicitly. c) Show that /epsilon1ijk/epsilon1ij/primek/prime=δjj/primeδkk/prime−δjk/primeδkj/prime. d) For dimension n= 4, write out /epsilon1ijkl/epsilon1ij/primek/primel/primeas a sum of products of δ’s similar to the one in part (c). Exercise 1.6 :Vector Products . The vector product of two three-vectors may be written in Cartesian components as ( a×b)i=/epsilon1ijkajbk. Use this and your results about /epsilon1ijkfrom the previous exercise to show that i)a·(b×c) =b·(c×a) =c·(a×b), ii)a×(b×c) = (a·c)b−(a·b)c, iii) (a×b)·(c×d) = (a·c)(b·d)−(a·d)(b·c). iv) If we take a,b,candd, with d≡b, to be unit vectors, show that the identities (i) and (iii) become the sine and cosine rule, respectively, of spherical trigonometry. (Hint: for the spherical sine ru le, begin by showing that a·[(a×b)×(a×c)] =a·(b×c).) 1.3.2 Stress and strain As an illustration of the utility of Cartesian tensors, we co nsider their appli- cation to elasticity. Suppose that an elastic body is slightly deformed so that the particle that was originally at the point with Cartesian co-ordinates xiis moved to xi+ηi. We define the (infinitesimal) strain tensor eijby eij=1 2/parenleftbigg∂ηj ∂xi+∂ηi ∂xj/parenrightbigg . (1.90) It is automatically symmetric: eij=eji. We will leave for later (exercise 2.3) a discussion of why this is the natural definition of stra in, and also the modifications necessary were we to employ a non-Cartesia n co-ordinate system. To define the stress tensor σijwe consider the portion Ω of the body in figure 1.1, and an element of area dS=nd|S|on its boundary. Here, nis 1.3. CARTESIAN TENSORS 21 the unit normal vector pointing out of Ω. The force Fexerted on this surface element by the parts of the body exterior to Ω has components Fi=σijnjd|S|. (1.91) Ω dF n |S| Figure 1.1: Stress forces. ThatFis a linear function of nd|S|can be seen by considering the forces on an small tetrahedron, three of whose sides coincide with t he co-ordinate planes, the fourth side having nas its normal. In the limit that the lengths of the sides go to zero as /epsilon1, the mass of the body scales to zero as /epsilon13, but the forces are proprtional to the areas of the sides and go to z ero only as /epsilon12. Only if the linear relation holds true can the acceleration o f the tetrahedron remain finite. A similar argument applied to torques and the m oment of intertia of a small cube shows that σij=σji. A generalization of Hooke’s law, σij=cijklekl, (1.92) relates the stress to the strain via the tensor of elastic constants cijkl. This rank-four tensor has the symmetry properties cijkl=cklij=cjikl=cijlk. (1.93) In other words, the tensor is symmetric under the interchang e of the first and second pairs of indices, and also under the interchange o f the individual indices in either pair. For an isotropic material — a material whose properties are i nvariant under the rotation group SO(3) — the tensor of elastic consta nts must be an 22 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE isotropic tensor. The most general such tensor with the requ ired symmetries is cijkl=λδijδkl+µ(δikδjl+δilδjk). (1.94) As isotropic material is therefore characterized by only tw o independent pa- rameters,λandµ. These are called the Lam´ e constants after the mathemat- ical engineer Gabriel Lam´ e. In terms of them the generalize d Hooke’s law becomes σij=λδijekk+ 2µeij. (1.95) By considering particular deformations, we can express the more directly measurable bulk modulus ,shear modulus ,Young’s modulus andPoisson’s ratioin terms of λandµ. The bulk modulus κis defined by dV V=−κdP, (1.96) where an infinitesimal isotropic external pressure dPcauses a change V→ V+dVin the volume of the material. This applied pressure corresp onds to a surface stress of σij=−δijdP. An isotropic expansion displaces points in the material so that ηi=1 3dV Vxi. (1.97) The strains are therefore given by eij=1 3δijdV V. (1.98) Inserting this strain into the stress-strain relation give s σij=δij(λ+2 3µ)dV V=−δijdP. (1.99) Thus κ=λ+2 3µ. (1.100) To define the shear modulus, we assume a deformation η1=θx2, so e12=e21=θ/2, with all other eijvanishing. 1.3. CARTESIAN TENSORS 23 σ21σ21 σ12σ12 θ Figure 1.2: Shear strain. The arrows show the direction of the applied stresses. The σ21on the vertical faces are necessary to stop the body ro- tating. The applied shear stress is σ12=σ21. The shear modulus, is defined to be σ12/θ. Inserting the strain components into the stress-strain re lation gives σ12=µθ, (1.101) and so the shear modulus is equal to the Lam´ e constant µ. We can therefore write the generalized Hooke’s law as σij= 2µ(eij−1 3δijekk) +κekkδij, (1.102) which reveals that the shear modulus is associated with the t raceless part of the strain tensor, and the bulk modulus with the trace. Young’s modulus Yis measured by stretching a wire of initial length L and square cross section of side Wunder a tension T=σ33W2. L σ33 σ 33W Figure 1.3: Forces on a stretched wire. We defineYso that σ33=YdL L. (1.103) At the same time as the wire stretches, its width changes W→W+dW. Poisson’s ratio σis defined by dW W=−σdL L, (1.104) 24 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE so thatσis positive if the wire gets thinner as it gets longer. The dis place- ments are η3=z/parenleftbiggdL L/parenrightbigg , η1=x/parenleftbiggdW W/parenrightbigg =−σx/parenleftbiggdL L/parenrightbigg , η2=y/parenleftbiggdW W/parenrightbigg =−σy/parenleftbiggdL L/parenrightbigg , (1.105) so the strain components are e33=dL L, e 11=e22=dW W=−σe33. (1.106) We therefore have σ33= (λ(1−2σ) + 2µ)/parenleftbiggdL L/parenrightbigg , (1.107) leading to Y=λ(1−2σ) + 2µ. (1.108) Now, the side of the wire is a free surface with no forces actin g on it, so 0 =σ22=σ11= (λ(1−2σ)−2σµ)/parenleftbiggdL L/parenrightbigg . (1.109) This tells us that6 σ=1 2λ λ+µ, (1.110) and Y=µ/parenleftbigg3λ+ 2µ λ+µ/parenrightbigg . (1.111) Other relations, following from those above, are Y= 3κ(1−2σ), = 2µ(1 +σ). (1.112) 6Poisson and Cauchy believed that λ=µ, and hence that σ= 1/4. 1.3. CARTESIAN TENSORS 25 Exercise 1.7 : Show that the symmetries cijkl=cklij=cjikl=cijlk imply that a general homogeneous material has 21 independen t elastic con- stants. (This result was originally obtained by George Gree n, of Green func- tion fame.) Exercise 1.8 : A steel beam is forged so that its cross section has the shape of a region Γ∈R2. When undeformed, it lies along the zaxis. The centroid O of each cross section is defined so that /integraldisplay Γxdxdy =/integraldisplay Γydxdy = 0, when the co-ordinates x,yare taken with the centroid O as the origin. The beam is slightly bent away from the zaxis so that the line of centroids remains in they,zplane. At a particular cross section with centroid O, the lin e of centroids has radius of curvature R. Γzxy O Figure 1.4: Bent beam. Assume that the deformation in the vicinity of O is such that ηx=−σ Rxy, ηy=1 2R/braceleftbig σ(x2−y2)−z2/bracerightbig , ηz=1 Ryz. 26 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE OΓ xy Figure 1.5: The original (dashed) and anticlastically deformed (full) cross- section. For positive Poisson ratio, the cross section deforms anticlastically — the sides bendupas the beam bends down. Compute the strain tensor resulting from the given deformat ion, and show that its only non-zero components are exx=−σ Ry, eyy=−σ Ry, ezz=1 Ry. Next, show that σzz=/parenleftbiggY R/parenrightbigg y, and that all other components of the stress tensor vanish. De duce from this vanishing that the assumed deformation satisfies the free-s urface boundary condition, and so is indeed the way the beam responds when it i s bent by forces applied at its ends. The work done in bending the beam /integraldisplay beam1 2eijcijklekld3x is stored as elastic energy. Show that for our bent rod this en ergy is equal to /integraldisplayYI 2/parenleftbigg1 R2/parenrightbigg ds≈/integraldisplayYI 2(y/prime/prime)2dz, wheresis the arc-length taken along the line of centroids of the bea m, I=/integraldisplay Γy2dxdy is the moment of inertia of the region Γ about the xaxis, andy/prime/primedenotes the second derivative of the deflection of the beam with respe ct toz(which 1.3. CARTESIAN TENSORS 27 approximates the arc-length). This last formula for the str ain energy has been used in a number of our calculus-of-variations problems. y z Figure 1.6: The distribution of forces σzzexerted on the left-hand part of the bent rod by the material to its right. 1.3.3 Maxwell stress tensor Consider a small cubical element of an elastic body. If the st ress tensor were position independent, the external forces on each pair of op posing faces of the cube would be equal in magnitude but pointing in opposite directions. There would therefore be no net external force on the cube. Wh enσijisnot constant then we claim that the total force acting on an infini tesimal element of volumedVis Fi=∂jσijdV. (1.113) To see that this assertion is correct, consider a finite regio n Ω with boundary ∂Ω, and use the divergence theorem to write the total force on Ω as Ftot i=/integraldisplay ∂Ωσijnjd|S|=/integraldisplay Ω∂jσijdV. (1.114) Whenever the force-per-unit-volume fiacting on a body can be written in the form fi=∂jσij, we refer to σijas a “stress tensor,” by analogy with stress in an elastic solid. As an example, let EandBbe electric and magnetic fields. For simplicity, initially assume them to be static. T he force per unit volume exerted by these fields on a distribution of charge ρand current jis f=ρE+j×B. (1.115) From Gauss’ law ρ= divD, and with D=/epsilon10E, we find that the force per unit volume due the electric field has components ρEi= (∂jDj)Ei=/epsilon10/parenleftBig ∂j(EiEj)−Ej∂jEi/parenrightBig 28 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE =/epsilon10/parenleftBig ∂j(EiEj)−Ej∂iEj/parenrightBig =/epsilon10∂j/parenleftbigg EiEj−1 2δij|E|2/parenrightbigg . (1.116) Here, in passing from the first line to the second, we have used the fact that curlEis zero for static fields, and so ∂jEi=∂iEj. Similarly, using j= curl H, together with B=µ0Hand div B= 0, we find that the force per unit volume due the magnetic field has components (j×B)i=µ0∂j/parenleftbigg HiHj−1 2δij|H|2/parenrightbigg . (1.117) The quantity σij=/epsilon10/parenleftbigg EiEj−1 2δij|E|2/parenrightbigg +µ0/parenleftbigg HiHj−1 2δij|H|2/parenrightbigg (1.118) is called the Maxwell stress tensor . Its utility lies in in the fact that the total electromagnetic force on an isolated body is the integ ral of the Maxwell stress over its surface. We do not need to know the fields withi n the body. Michael Faraday was the first to intuit a picture of electroma gnetic stresses and attributed both a longitudinal tension and a mutual late ral repulsion to the field lines. Maxwell’s tensor expresses this idea mathem atically. Exercise 1.9 : Allow the fields in the preceding calculation to be time depe n- dent. Show that Maxwell’s equations curlE=−∂B ∂t,divB= 0, curlH=j+∂D ∂t,divD=ρ, withB=µ0H,D=/epsilon10E, andc= 1/√µ0/epsilon10, lead to (ρE+j×B)i+∂ ∂t/braceleftbigg1 c2(E×H)i/bracerightbigg =∂jσij. The left-hand side is the time rate of change of the mechanica l (first term) and electromagnetic (second term) momentum density. Obser ve that we can equivalently write ∂ ∂t/braceleftbigg1 c2(E×H)i/bracerightbigg +∂j(−σij) =−(ρE+j×B)i, 1.4. FURTHER EXERCISES AND PROBLEMS 29 and think of this a local field-momentum conservation law. In this interpre- tation−σijis thought of as the momentum flux tensor, its entries being the flux in direction jof the component of field momentum in direction i. The term on the right-hand side is the rate at which momentum is be ing supplied to the electro-magnetic field by the charges and currents. 1.4 Further Exercises and Problems Exercise 1.10 :Quotient theorem. Suppose that you have come up with some recipe for generating an array of numbers Tijkin any co-ordinate frame, and want to know whether these numbers are the components of a tri ply con- travariant tensor. Suppose further that you know that, give n the components aijof an arbitrary doubly covariant tensor, the numbers Tijkajk=vi transform as the components of a contravariant vector. Show thatTijkdoes indeed transform as a triply contravariant tensor. (The nat ural generalization of this result to arbitrary tensor types is known as the quotient theorem .) Exercise 1.11 : LetTijbe the 3-by-3 array of components of a tensor. Show that the quantities a=Tii, b=TijTji, c=TijTjkTki are invariant. Further show that the eigenvalues of the line ar map represented by the matrix Tijcan be found by solving the cubic equation λ3−aλ2+1 2(a2−b)λ−1 6(a3−3ab+ 2c) = 0. Exercise 1.12 : Let the covariant tensor Rijklpossess the following symme- tries: i)Rijkl=−Rjikl, ii)Rijkl=−Rijlk, iii)Rijkl+Riklj+Riljk= 0. Use the properties i),ii), iii) to show that: a)Rijkl=Rklij. b) IfRijklxiyjxkyl= 0 for all vectors xi,yi, thenRijkl= 0. 30 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE c) IfBijis a symmetric covariant tensor and set we Aijkl=BikBjl−BilBjk, thenAijklhas the same symmetries as Rijkl. Exercise 1.13 : Write out Euler’s equation for fluid motion ˙v+ (v·∇)v=−∇h in Cartesian tensor notation. Transform it into ˙v−v×ω=−∇/parenleftbigg1 2v2+h/parenrightbigg , whereω=∇×vis the vorticity. Deduce Bernoulli’s theorem, that for stea dy (˙v= 0) flow the quantity1 2v2+his constant along streamlines. Exercise 1.14 :Symmetric integration . Show that the n-dimensional integral Iαβγδ=/integraldisplaydnk (2π)n(kαkβkγkδ)f(k2), is equal to A(δαβδγδ+δαγδβδ+δαδδβγ) where A=1 n(n+ 2)/integraldisplaydnk (2π)n(k2)2f(k2). Similarly evaluate Iαβγδ/epsilon1=/integraldisplaydnk (2π)n(kαkβkγkδk/epsilon1)f(k2). Exercise 1.15 : Write down the most general three-dimensional isotropic t en- sors of rank two and three. In piezoelectric materials, the application of an electric fieldEiinduces a mechanical strain that is described by a rank-two symmetric tensor eij=dijkEk, wheredijkis a third-rank tensor that depends only on the material. Sho w thateijcan only be non-zero in an anisotropic material. 1.4. FURTHER EXERCISES AND PROBLEMS 31 Exercise 1.16 : In three dimensions, a rank-five isotropic tensor Tijklmis a linear combination of expressions of the form /epsilon1i1i2i3δi4i5for some assignment of the indices i,j,k,l,m to thei1,...,i 5. Show that, on taking into account the symmetries of the Kronecker and Levi-Civita symbols, we can construct tendistinct products /epsilon1i1i2i3δi4i5. Only sixof these are linearly independent, however. Show, for example, that /epsilon1ijkδlm−/epsilon1jklδim+/epsilon1kliδjm−/epsilon1lijδkm= 0, and find the three other independent relations of this sort.7 (Hint: Begin by showing that, in three dimensions, δi1i2i3i4 i5i6i7i8def=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleδi1i5δi1i6δi1i7δi1i8 δi2i5δi2i6δi2i7δi2i8 δi3i5δi3i6δi3i7δi3i8 δi4i5δi4i6δi4i7δi4i8/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0, and contract with /epsilon1i6i7i8.) Problem 1.17 :The Pl¨ ucker Relations. This problem provides a challenging test of your understanding of linear algebra. It leads you th rough the task of deriving the necessary and sufficient conditions for A=Ai1...ikei1∧...∧eik∈/logicalanddisplaykV to be decomposable as A=f1∧f2∧...∧fk. The trick is to introduce two subspaces of V, i)W, the smallest subspace of Vsuch that A∈/logicalandtextkW, ii)W/prime={v∈V:v∧A= 0}, and explore their relationship. a) Show that if{w1,w2,...,wn}constitute a basis for W/prime, then A=w1∧w2∧···∧ wn∧ϕ for someϕ∈/logicalandtextk−nV. Conclude that that W/prime⊆W, and that equal- ity holds if and only if Ais decomposable, in which case W=W/prime= span{f1...fk}. 7Such relations are called syzygies . A recipe for constructing linearly independent basis sets of isotropic tensors can be found in: G. F. Smith, Tensor ,19(1968) 79-88. 32 CHAPTER 1. TENSORS IN EUCLIDEAN SPACE b) Now show that Wis the image space of/logicalandtextk−1V∗under the map that takes Ξ= Ξi1...ik−1e∗i1∧...∧e∗ik−1∈/logicalanddisplayk−1V∗ to i(Ξ)Adef= Ξi1...ik−1Ai1...ik−1jej∈V Deduce that the condition W⊆W/primeis that /parenleftBig i(Ξ)A/parenrightBig ∧A= 0,∀Ξ∈/logicalanddisplayk−1V∗. c) By taking Ξ=e∗i1∧...∧e∗ik−1, show that the condition in part b) can be written as Ai1...ik−1j1Aj2j3...jk+1ej1∧...∧ejk+1= 0. Deduce that the necessary and sufficient conditions for decom posibility are that Ai1...ik−1[j1Aj2j3...jk+1]= 0, for all possible index sets i1,...,ik−1,j1,...jk+1. Here [...] denotes anti- symmetrization of the enclosed indices. Chapter 2 Differential Calculus on Manifolds In this section we will apply what we have learned about vecto rs and tensors in a linear space to the case of vector and tensor fieldsin a general curvilinear co-ordinate system. Our aim is to introduce the reader to the modern lan- guage of advanced calculus, and in particular to the calculu s of differential forms on surfaces and manifolds. 2.1 Vector and Covector Fields Vector fields — electric, magnetic, velocity fields, and so on — appear every- where in physics. After perhaps struggling with it in introd uctory courses, we rather take the field concept for granted. There remain subtl eties, however. Consider an electric field. It makes sense to add two field vect ors at a single point, but there is no physical meaning to the sum of field vect orsE(x1) and E(x2) at two distinct points. We should therefore regard all poss ible electric fields at a single point as living in a vector space, but each di fferent point in space comes with its own field-vector space. This view seem s even more reasonable when we consider velocity vectors describing mo tion on a curved surface. A velocity vector lives in the tangent space to the surface at each point, and each of these spaces is a differently oriented subspace of the higher- dimensional ambient space. 33 34 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS Figure 2.1: Each point on a surface has its own vector space of tangents. Mathematicians call such a collection of vector spaces — one for each of the points in a surface — a vector bundle over the surface. Thus the tangent bundle over a surface is the totality of all vector spaces tangent to the surface. Why a bundle ? This word is used because the individual tangent spaces are not completely independent, but are tied together in a rathe r non-obvious way. Try to construct a smooth field of unit vectors tangent to the surface of a sphere. However hard you work you will end up in trouble so mewhere. You cannot comb a hairy ball. On the surface of torus you will h ave no problems. You can comb a hairy doughnut. The tangent spaces c ollectively know something about the surface they are tangent to. Although we spoke in the previous paragraph of vectors tange nt to a curved surface, it is useful to generalize this idea to vecto rs lying in the tangent space of an n-dimensional manifold . Ann-manifoldMis essentially a space that locally looks like a part of Rn. This means that some open neighbourhood of each point can be parametrized by an n-dimensional co- ordinate system. Such a parametrization is called a chart. UnlessMisRn itself (or part of it), a chart will cover only part of M, and more than one will be required for complete coverage. Where a pair of chart s overlap we demand that the transformation formula giving one set of co- ordinates as a function of the other be a smooth ( C∞) function, and to possess a smooth inverse.1A collection of such smoothly related co-ordinate charts co vering all ofMis called an atlas. The advantage of thinking in terms of manifolds is that we do not have to understand their properties as arisi ng from some embedding in a higher dimensional space. Whatever structur e they have, they possess in, and of, themselves 1A formal definition of a manifold contains some further techn ical restrictions (that the space be Hausdorff andparacompact ) that are designed to eliminate pathologies. We are more interested in doing calculus than in proving theorems, and so we will ignore these niceties. 2.1. VECTOR AND COVECTOR FIELDS 35 Classical mechanics provides a familiar illustration of th ese ideas. The configuration space Mof a mechanical system is usually a manifold. When the system has ndegrees of freedom we use generalized co-ordinates qi,i= 1,...,n to parameterize M. The tangent bundle of Mthen provides the setting for Lagrangian mechanics. This bundle, denoted by TM, is the 2n- dimensional space whose points consist of a point pinMtogether with a tangent vector lying in the tangent space TMpat that point. If we think of the tangent vector as a velocity, the natural co-ordinate s onTMbecome (q1,q2,...,qn; ˙q1,˙q2,...,˙qn), and these are the variables that appear in the Lagrangian of the system. If we consider a vector tangent to some curved surface, it wil l stick out of it. If we have a vector tangent to a manifold, it is a straigh t arrow lying atop bent co-ordinates. Should we restrict the length of the vector so that it does not stick out too far? Are we restricted to only infinit esimal vectors? It’s best to avoid all this by inventing a clever notion of wha t a vector in a tangent space is. The idea is to focus on a well-defined objec t such as a derivative. Suppose our space has co-ordinates xµ(These are notthe contravariant components of some vector). A directional derivative is an object such as Xµ∂µwhere∂µis shorthand for ∂/∂xµ. When the numbers Xµare functions of the co-ordinates xσ, this object is called a tangent-vector field, and we write2 X=Xµ∂µ. (2.1) We regard the ∂µat a pointxas a basis for TMx, the tangent-vector space at x, and theXµ(x) as the (contravariant) components of the vector Xat that point. Although they are not little arrows, what the ∂µare is mathematically clear, and so we know perfectly well how to deal with them. When we change co-ordinate system from xµtozνby regarding the xµ’s as invertable functions of the zν’s,i.e. x1=x1(z1,z2,...,zn), x2=x2(z1,z2,...,zn), ... xn=xn(z1,z2,...,zn), (2.2) 2We are going to stop using bold symbols to distinguish betwee n intrinsic objects and their components, because from now on almost everything wil l be something other than a number, and too much black ink would just be confusing. 36 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS then the chain rule for partial differentiation gives ∂µ≡∂ ∂xµ=∂zν ∂xµ∂ ∂zν=/parenleftbigg∂zν ∂xµ/parenrightbigg ∂/prime ν, (2.3) where∂/prime νis shorthand for ∂/∂zν. By demanding that X=Xµ∂µ=X/primeν∂/prime ν (2.4) we find the components in the zνco-ordinate frame to be X/primeν=/parenleftbigg∂zν ∂xµ/parenrightbigg Xµ. (2.5) Conversely, using ∂xσ ∂zν∂zν ∂xµ=∂xσ ∂xν=δσ µ, (2.6) we have Xν=/parenleftbigg∂xν ∂zµ/parenrightbigg X/primeµ. (2.7) This, then, is the transformation law for a contravariant ve ctor. It is worth pointing out that the basis vectors ∂µarenotunit vectors. As we have no metric, and therefore no notion of length anyway, w e cannot try to normalize them. If you insist on drawing (small?) arrows, think of∂1as starting at a point ( x1,x2,...,xn) and with its head at ( x1+ 1,x2,...,xn). Of course this is only a good picture if the co-ordinates are n ot too “curvy.” x =2 x =3x =4 x =5 x =4x =6111 222 2 1 Figure 2.2: Approximate picture of the vectors ∂1and∂2at the point (x1,x2) = (2,4). Example: The surface of the unit sphere is a manifold. It is usually den oted byS2. We may label its points with spherical polar co-ordinates θandφ, 2.1. VECTOR AND COVECTOR FIELDS 37 and these will be useful everywhere except at the north and so uth poles, where they become singular because at θ= 0 orπall values of φcorrespond to the same point. In this co-ordinate basis, the tangent vec tor representing the velocity field due to a rigid rotation of one radian per sec ond about the zaxis is Vz=∂φ. (2.8) Similarly Vx=−sinφ∂θ−cotθcosφ∂φ, Vy= cosφ∂θ−cotθsinφ∂φ, (2.9) represent rigid rotations about the xandyaxes. We now know how to think about vectors. What about their dual- space partners, the covectors? These live in the cotangent bundle T∗M, and for them a cute notational game, due to ´Elie Cartan, is played. We write the basis vectors dual to the ∂µasdxµ( ). Thus dxµ(∂ν) =δµ ν. (2.10) When evaluated on a vector field X=Xµ∂µ, the basis covectors dxµreturn its components dxµ(X) =dxµ(Xν∂ν) =Xνdxµ(∂ν) =Xνδµ ν=Xµ. (2.11) Now, any smooth function f∈C∞(M) will give rise to a field of covectors inT∗M. This is because a vector field Xacts on the scalar function fas Xf=Xµ∂µf (2.12) andXfis another scalar function. This new function gives a number — and thus an element of the field R— at each point x∈M. But this is exactly what a covector does: it takes in a vector at a point and return s a number. We will call this covector field “ df.” It is essentially the gradient of f. Thus df(X)def=Xf=Xµ∂f ∂xµ. (2.13) If we takefto be the co-ordinate xν, we have dxν(X) =Xµ∂xν ∂xµ=Xµδν µ=Xν, (2.14) 38 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS so this viewpoint is consistent with our previous definition ofdxν. Thus df(X) =∂f ∂xµXµ=∂f ∂xµdxµ(X) (2.15) for any vector field X. In other words, we can expand dfas df=∂f ∂xµdxµ. (2.16) This is notsome approximation to a change in f, but is an exact expansion of the covector field dfin terms of the basis covectors dxµ. We may retain something of the notion that dxµrepresents the (con- travariant) components of a small displacement in xprovided that we think ofdxµas a machine into which we insert the small displacement (a ve ctor) and have it spit out the numerical components δxµ. This is the same dis- tinction that we make between sin( ) as a function into which o ne can plug x, and sinx, the number that results from inserting in this particular v alue ofx. Although seemingly innocent, we know that it is a distincti on of great power. The change of co-ordinates transformation law for a covecto r fieldfµis found from fµdxµ=f/prime νdzν, (2.17) by using dxµ=/parenleftbigg∂xµ ∂zν/parenrightbigg dzν. (2.18) We find f/prime ν=/parenleftbigg∂xµ ∂zν/parenrightbigg fµ. (2.19) A general tensor such as Qλµ ρστtransforms as Q/primeλµ ρστ(z) =∂zλ ∂xα∂zµ ∂xβ∂xγ ∂zρ∂xδ ∂zσ∂x/epsilon1 ∂zτQαβ γδ/epsilon1(x). (2.20) Observe how the indices are wired up: Those for the new tensor coefficients in the new co-ordinates, z, are attached to the new z’s, and those for the old coefficients are attached to the old x’s. Upstairs indices go in the numerator of each partial derivative, and downstairs ones are in the de nominator. 2.2. DIFFERENTIATING TENSORS 39 The language of bundles and sections At the beginning of this section, we introduced the notion of a vector bundle. This is a particular example of the more general concept of a fibre bundle , where the vector space at each point in the manifold is replac ed by a “fibre” overthat point. The fibre can be any mathematical object, such as a set, tensor space, or another manifold. Mathematicians visuali ze the bundle as a collection of fibres growing out of the manifold, much as sta lks of wheat grow out the soil. When one slices through a patch of wheat wit h a scythe, the blade exposes a cross-section of the stalks. By analogy, a choice of an element of the the fibre over each point in the manifold is call ed across- section , or, more commonly, a section of the bundle. In this language a tangent-vector field becomes a section of the tangent bundle , and a field of covectors becomes a section of the cotangent bundle. We provide a more detailed account of bundles in chapter 7. 2.2 Differentiating Tensors Iffis a function then ∂µfare components of the covariant vector df. Suppose thataµis a contravariant vector. Are ∂νaµthe components of a type (1 ,1) tensor? The answer is no! In general, differentiating the components of a tensor does not give rise to another tensor. One can see why at two levels: a) Consider the transformation laws. They contain expressi ons of the form ∂xµ/∂zν. If we differentiate both sides of the transformation law of a tensor, these factors are also differentiated, but tensor tr ansformation laws never contain second derivatives, such as ∂2xµ/∂zν∂zσ. b) Differentiation requires subtracting vectors or tensors at different points — but vectors at different points are in different vector space s, so their difference is not defined. These two reasons are really one and the same. We need to be cle verer to get new tensors by differentiating old ones. 2.2.1 Lie Bracket One way to proceed is to note that the vector field Xis anoperator . It makes sense, therefore, to try to compose two of them to make anothe r. Look at 40 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS XY, for example: XY=Xµ∂µ(Yν∂ν) =XµYν∂2 µν+Xµ/parenleftbigg∂Yν ∂xµ/parenrightbigg ∂ν. (2.21) What are we to make of this? Not much! There is no particular in terpretation for the second derivative, and as we saw above, it does not tra nsform nicely. But suppose we take a commutator : [X,Y] =XY−YX= (Xµ(∂µYν)−Yµ(∂µXν))∂ν. (2.22) The second derivatives have cancelled, and what remains is a directional derivative and so a bona-fide vector field. The components [X,Y]ν≡Xµ(∂µYν)−Yµ(∂µXν) (2.23) arethe components of a new contravariant vector field made from t he two old vector fields. It is called the Lie bracket of the two fields, and has a geometric interpretation. To understand the geometry of the Lie bracket, we first define t heflow associated with a tangent-vector field X. This is the map that takes a point x0and maps it to x(t) by solving the family of equations dxµ dt=Xµ(x1,x2,...,xd), (2.24) with initial condition xµ(0) =xµ 0. In words, we regard Xas the velocity field of a flowing fluid, and let xride along with the fluid. Now envisage XandYas two velocity fields. Suppose we flow along X for a brief time t, then along Yfor another brief interval s. Next we switch back toX, but with a minus sign, for time t, and then to−Yfor a final interval ofs. We have tried to retrace our path, but a short exercise with Taylor’s theorem shows that we will fail to return to our exac t starting point. We will miss by δxµ=st[X,Y]µ, plus corrections of cubic order in sandt. 2.2. DIFFERENTIATING TENSORS 41 −sYtXsY −tX X,Y[ ]st Figure 2.3: The Lie bracket. Example: Let Vx=−sinφ∂θ−cotθcosφ∂φ, Vy= cosφ∂θ−cotθsinφ∂φ be two vector fields in T(S2). We find that [Vx,Vy] =−Vz, whereVz=∂φ. Frobenius’ Theorem Suppose that in some region of a d-dimensional manifold Mwe are given n < d linearly independent tangent-vector fields Xi. Such a set is called a distribution by differential geometers. (The concept has nothing to do wit h probability, or with objects like “ δ(x)” which are also called “distributions.”) At each point x, the span/angbracketleftXi(x)/angbracketrightof the field vectors vectors forms a subspace of the tangent space TMx, and we can picture this subspace as a fragment of ann-dimensional surface passing through x. It is possible that these surface fragments fit together to make a stack of smooth surfa ces — called a foliation — that fill out the d-dimensional space, and have the given Xias their tangent vectors. 42 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS X1X2 xN Figure 2.4: A local foliation. If this is the case then starting from xand taking steps only along the Xi we find ourselves restricted to the n-surface, or n-submanifold ,Npassing though the original point x. Alternatively, the surface fragments may form such an incoh erent jumble that starting from xand moving only along the Xiwe can find our way to any point in the neighbourhood of x. It is also possible that some intermediate case applies, so that moving along the Xirestricts us to an m-surface, where d > m > n . The Lie bracket provides us with the appropriate tool with which to investigate these possibilities. First a definition: If there are functions ck ij(x) such that [Xi,Xj] =ck ij(x)Xk, (2.25) i.e.the Lie brackets close within the set {Xi}at each point x, then the distribution is said to be involutive. When our given distribution is involutive, then the first case holds, and, at least locally, there is a fol iation byn- submanifolds N. A formal statement of this is: Theorem (Frobenius): A smooth ( C∞) involutive distribution is completely integrable : locally, there are co-ordinates xµ,µ= 1,...,d such thatXi=/summationtextn µ=1Xµ i∂µ, and the surfaces Nthrough each point are in the form xµ= const. forµ=n+ 1,...,d . Conversely, if such co-ordinates exist then the distribution is involutive. Sketch of Proof : If such co-ordinates exist then it is obvious that the Lie bracket of any pair of vectors in the form Xi=/summationtextn µ=1Xµ i∂µcan also be ex- panded in terms of the first nbasis vectors. A logically equivalent statement exploits the geometric interpretation of the Lie bracket: I f the Lie brackets of the fields Xidonotclose within the n-dimensional span of the Xi, then a sequence of back-and-forth manouvres along the Xiallows us to escape into a new direction, and so the Xicannot be tangent to an n-surface. Establishing 2.2. DIFFERENTIATING TENSORS 43 the converse — that closure implies the existence of the foli ation — is rather more technical, and we will not attempt it. The physicist’s version of Frobenius’ theorem is usually ex pressed in the language of holonomic oranholonomic constraints. For example, consider a particle moving in three dimensions . If we are told that the velocity vector is constrained to be perpendic ular to the radius vector, i.e.v·r= 0, we realize that the particle is being forced to move on a the sphere|r|=r0passing through the initial point. In spherical co-ordinat es the associated distribution is the set {∂θ,∂φ}, which is clearly involutive. The foliation is the family of nested spheres whose centre is the origin. The foliation is not global because it becomes singular at r= 0. Constraints like this, which restrict the motion to a surface, are called holonomic . Suppose, on the other hand, we have a ball rolling on a table. H ere, we have a five-dimensional configuration manifold M=R2×S3parameterized by the centre of mass ( x,y)∈R2of the ball and the three Euler angles (θ,φ,ψ )∈S3defining its orientation. Three no-slip rolling conditions ˙x= ˙ψsinθsinφ+˙θcosφ, ˙y=−˙ψsinθcosφ+˙θsinφ, 0 = ˙ψcosθ+˙φ, (2.26) (see exercise 2.17) link the rate of change of the Euler angle s to the velocity of the centre of mass. At each point in this five-dimensional m anifold we are free to roll the ball in two directions, and so might expect th at the reachable configurations constitute a two-dimensional surface embed ded in the full five- dimensional space. The two vector fields rollx=∂x−sinφcotθ∂φ+ cosφ∂θ+ cosecθsinφ∂ψ, rolly=∂y+ cosφcotθ∂φ+ sinφ∂θ−cosecθcosφ∂ψ,(2.27) describing the x- andy-direction rolling motion are not in involution, how- ever. By calculating enough Lie brackets we eventually obta in five linearly independent velocity vector fields, and starting from one co nfiguration we can reach any other. The no-slip rolling condition is said to be non-integrable , or anholonomic . Such systems are tricky to deal with in Lagrangian dynamics . For ad-dimensional mechanical system, a set of mindependent con- straints of the form ωi µ(q) ˙qµ= 0,i= 1,...,m determines an n=d−m 44 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS dimensional distribution. In terms of the vector ˙ q≡˙qµ∂µand the covectors ωi=d/summationdisplay µ=1ωi µ(q)dqµ, i= 1≤i≤m (2.28) we can write the these constraints as ωi( ˙q) = 0. This is known a Pfaffian system of equations. The Pfaffian system is said to be integrable if the distribution it implicitly defines is in involution, and hen ce itself integrable. In this case there is a set of mfunctionsgi(q) and an invertible m-by-m matrixfi j(q) such that ωi=m/summationdisplay j=1fi j(q)dgj. (2.29) The functions gi(q) can, for example, be taken to be the co-ordinate func- tionsxµ,µ=n+ 1,...,d , that label the foliating surfaces Nin the state- ment of Frobenius’ theorem. The system of integrable constr aintsωi( ˙q) = 0 thus restricts us to the surfaces gi(q) =constant . Integrable constraints are therefore holonomic. The following exercise provides a familiar example of the ut ility of non- holonomic constraints: Exercise 2.1 :Parallel Parking using Lie Brackets . θ (x,y)drive parkφ Figure 2.5: Co-ordinates for car parking 2.2. DIFFERENTIATING TENSORS 45 The configuration space of a car is four dimensional, and para meterized by co-ordinates ( x,y,θ,φ ), as shown in figure 2.5. Define the following vector fields: a) (front wheel) drive = cosφ(cosθ∂x+ sinθ∂y) + sinφ∂θ. b)steer =∂φ. c) (front wheel) skid=−sinφ(cosθ∂x+ sinθ∂y) + cosφ∂θ. d)park =−sinθ∂x+ cosθ∂y. Explain why these are apt names for the vector fields, and comp ute the Lie brackets: [steer,drive ],[steer,skid],[skid,drive ], [park,drive ],[park,park],[park,skid]. The driver can use only the operations ( ±)drive and (±)steer to manouvre the car. Use the geometric interpretation of the Lie bracket to explain how a suitable sequence of motions (forward, reverse, and turnin g the steering wheel) can be used to manoeuvre a car sideways into a parking space. 2.2.2 Lie Derivative Another derivative we can define is the Lie derivative along a vector field X. It is defined by its action on a scalar function fas LXfdef=Xf, (2.30) on a vector field by LXYdef= [X,Y], (2.31) and on anything else by requiring it to be a derivation , meaning that it obeys Leibniz’ rule. For example, let us compute the Lie derivativ e of a covector F. We first introduce an arbitrary vector field Yand plug it into Fto get the scalar function F(Y). Leibniz’ rule is then the statement that LXF(Y) = (LXF)(Y) +F(LXY). (2.32) SinceF(Y) is a function and Ya vector, both of whose derivatives we know how to compute, we know two of the three terms in this equation . From LXF(Y) =XF(Y) andF(LXY) =F([X,Y]), we have XF(Y) = (LXF)(Y) +F([X,Y]), (2.33) 46 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS and so (LXF)(Y) =XF(Y)−F([X,Y]). (2.34) In components, this becomes (LXF)(Y) =Xν∂ν(FµYµ)−Fν(Xµ∂µYν−Yµ∂µXν) = (Xν∂νFµ+Fν∂µXν)Yµ. (2.35) Note how all the derivatives of Yµhave cancelled, so LXF( ) depends only on the local value of Y. The Lie derivative of Fis therefore still a covector field. This is true in general: the Lie derivative does not cha nge the tensor character of the objects on which it acts. Dropping the passi ve spectator fieldYν, we have a formula for LXFin components: (LXF)µ=Xν∂νFµ+Fν∂µXν. (2.36) Another example is provided by the Lie derivative of a type (0 ,2) tensor, such as a metric tensor. This is (LXg)µν=Xα∂αgµν+gµα∂νXα+gαν∂µXα. (2.37) The Lie derivative of a metric measures the extent to which th e displacement xα→xα+/epsilon1Xα(x) deforms the geometry. If we write the metric as g(,) =gµν(x)dxµ⊗dxν, (2.38) we can understand both this geometric interpretation and th e origin of the three terms appearing in the Lie derivative. We simply make t he displace- mentxα→xα+/epsilon1Xαin the coefficients gµν(x) and in the two dxα. In the latter we write d(xα+/epsilon1Xα) =dxα+/epsilon1∂Xα ∂xβdxβ. (2.39) Then we see that gµν(x)dxµ⊗dxν→[gµν(x) +/epsilon1(Xα∂αgµν+gµα∂νXα+gαν∂µXα)]dxµ⊗dxν = [gµν+/epsilon1(LXg)µν]dxµ⊗dxν. (2.40) A displacement field Xthat does not change distances between points, i.e. one that gives rise to an isometry , must therefore satisfy LXg= 0. Such an Xis said to be a Killing field after Wilhelm Killing who introduced them in his study of non-euclidean geometries. 2.2. DIFFERENTIATING TENSORS 47 The geometric interpretation of the Lie derivative of a vect or field is as follows: In order to compute the Xdirectional derivative of a vector field Y, we need to be able to subtract the vector Y(x) from the vector Y(x+/epsilon1X), divide by/epsilon1, and take the limit /epsilon1→0. To do this we have somehow to get the vectorY(x) from the point x, where it normally resides, to the new point x+/epsilon1X, so both vectors are elements of the same vector space. The Li e derivative achieves this by carrying the old vector to the ne w point along the fieldX. Xε xLε XεYX Y(x+εX) Y(x) Figure 2.6: Computing the Lie derivative of a vector. Imagine the vector Yas drawn in ink in a flowing fluid whose velocity field isX. Initially the tail of Yis atxand its head is at x+Y. After flowing for a time/epsilon1, its tail is at x+/epsilon1X—i.eexactly where the tail of Y(x+/epsilon1X) lies. Where the head of transported vector ends up depends ho w the flow has stretched and rotated the ink, but it is this distorted vecto r that is subtracted fromY(x+/epsilon1X) to get/epsilon1LXY=/epsilon1[X,Y]. Exercise 2.2 : The metric on the unit sphere equipped with polar co-ordina tes is g(,) =dθ⊗dθ+ sin2θdφ⊗dφ. Consider Vx=−sinφ∂θ−cotθcosφ∂φ, the vector field of a rigid rotation about the xaxis. Compute the Lie derivative LVxg, and show that it is zero. Exercise 2.3 : Suppose we have an unstrained block of material in real spac e. A co-ordinate system ξ1,ξ2,ξ3, is attached to the atoms of the body. The point with co-ordinate ξis located at ( x1(ξ),x2(ξ),x3(ξ)) wherex1,x2,x3are the usual R3Cartesian co-ordinates. 48 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS a) Show that the induced metric in the ξco-ordinate system is gµν(ξ) =3/summationdisplay a=1∂xa ∂ξµ∂xa ∂ξν. b) The body is now deformed by an infinitesimal strain vector fi eldη(ξ). The atom with co-ordinate ξµis moved to what was ξµ+ηµ(ξ), or equiv- alently, the atom initially at Cartesian co-ordinate xa(ξ) is moved to xa+ηµ∂xa/∂ξµ. Show that the new induced metric is gµν+δgµν=gµν+Lηgµν. c) Define the strain tensor to be 1/2 of the Lie derivative of the metric with respect to the deformation. If the original ξco-ordinate system coincided with the Cartesian one, show that this definition r educes to the familiar form eab=1 2/parenleftbigg∂ηa ∂xb+∂ηb ∂xa/parenrightbigg , all tensors being Cartesian. d) Part c) gave us the geometric definitition of infinitesimal strain . If the body is deformed substantially, the Cauchy-Green finite strain tensor is defined as Eµν(ξ) =1 2/parenleftBig gµν−g(0) µν/parenrightBig , whereg(0) µνis the metric in the undeformed body and gµνthat of the deformed body. Explain why this is a reasonable definition. 2.3 Exterior Calculus 2.3.1 Differential Forms The objects we introduced in section 2.1, the dxµ, are called one-forms, or differential one-forms. They are fields living in the cotange nt bundleT∗M ofM. More precisely, they are sections of the cotangent bundle. Sections of the bundle whose fibre above x∈Mis thep-th skew-symmetric tensor power/logicalandtextp(T∗Mx) of the cotangent space are known as p-forms. For example, A=Aµdxµ=A1dx1+A2dx2+A3dx3, (2.41) 2.3. EXTERIOR CALCULUS 49 is a 1-form, F=1 2Fµνdxµ∧dxν=F12dx1∧dx2+F23dx2∧dx3+F31dx3∧dx1,(2.42) is a 2-form, and Ω =1 3!Ωµνσdxµ∧dxν∧dxσ = Ω 123dx1∧dx2∧dx3, (2.43) is a 3-form. All the coefficients are skew-symmetric tensors, so, for example, Ωµνσ= Ωνσµ= Ωσµν=−Ωνµσ=−Ωµσν=−Ωσνµ. (2.44) In each example we have explicitly written out all the indepe ndent terms for the case of three dimensions. Note how the p! disappears when we do this and keep only distinct components. In ddimensions the space of p-forms is d!/p!(d−p)! dimensional, and all p-forms with p>d vanish identically. As with the wedge products in chapter one, we regard a p-form as a p- linear skew-symetric function with pslots into which we can drop vectors to get a number. For example the basis two-forms give dxµ∧dxν(∂α,∂β) =δµ αδν β−δµ βδν α. (2.45) The analogous expression for a p-form would have p! terms. We can define an algebra of differential forms by “wedging” them together i n the obvious way, so that the product of a pform with a qform is a (p+q)-form. The wedge product is associative and distributive but not, of co urse, commuta- tive. Instead, if ais ap-form andbaq-form, then a∧b= (−1)pqb∧a. (2.46) Actually it is customary in this game to suppress the “ ∧” and simply write F=1 2Fµνdxµdxν, it being assumed that you know that dxµdxν=−dxνdxµ — what else could it be? 2.3.2 The Exterior Derivative Thesep-forms may seem rather complicated, so it is perhaps surpris ing that all the vector calculus (div, grad, curl, the divergence the orem and Stokes’ 50 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS theorem, etc.) that you have learned in the past reduce, in terms of them, to two simple formulæ! Indeed ´Elie Cartan’s calculus of p-forms is slowly supplanting traditional vector calculus, much as Willard G ibbs’ and Oliver Heaviside’s vector calculus supplanted the tedious compon ent-by-component formulæ you find in Maxwell’s Treatise on Electricity and Magnetism . The basic tool is the exterior derivative “d”, which we now define ax- iomatically: i) Iffis a function (0-form), then dfcoincides with the previous defini- tion,i.e.df(X) =Xffor any vector field X. ii)dis ananti-derivation : Ifais ap-form andbaq-form then d(a∧b) =da∧b+ (−1)pa∧db. (2.47) iii)Poincar´ e’s lemma :d2= 0, meaning that d(da) = 0 for any p-forma. iv)dis linear. That d(αa) =αda, for constant αfollows already from i) and ii), so the new fact is that d(a+b) =da+db. It is not immediately obvious that axioms i), ii) and iii) are compatible with one another. If we use axiom i), ii) and d(dxi) = 0 to compute the dof Ω =1 p!Ωi1,...,ipdxi1···dxip, we find dΩ =1 p!(dΩi1,...,ip)dxi1···dxip =1 p!∂kΩi1,...,ipdxkdxi1···dxip. (2.48) Now compute d(dΩ) =1 p!/parenleftbig ∂l∂kΩi1,...,ip/parenrightbig dxldxkdxi1···dxip. (2.49) Fortunately this is zero because ∂l∂kΩ =∂k∂lΩ, whiledxldxk=−dxkdxl. IfA=A1dx1+A2dx2+A3dx3then dA=/parenleftbigg∂A2 ∂x1−∂A1 ∂x2/parenrightbigg dx1dx2+/parenleftbigg∂A1 ∂x3−∂A3 ∂x1/parenrightbigg dx3dx1+/parenleftbigg∂A3 ∂x2−∂A2 ∂x3/parenrightbigg dx2dx3 =1 2Fµνdxµdxν, (2.50) where Fµν≡∂µAν−∂νAµ. (2.51) 2.3. EXTERIOR CALCULUS 51 You will recognize the components of curl Ahiding in here. Similarly, if F=F12dx1dx2+F23dx2dx3+F31dx3dx1then dF=/parenleftbigg∂F23 ∂x1+∂F31 ∂x2+∂F12 ∂x3/parenrightbigg dx1dx2dx3. (2.52) This looks like a divergence. The axiom d2= 0 encompasses both “curlgrad = 0” and “div curl = 0”, together with an infinite number of higher-dimensional a nalogues. The familiar “curl =∇×”, meanwhile, is only defined in three dimensional space. The exterior derivative takes p-forms to (p+1)-forms i.e.skew-symmetric type (0,p) tensors to skew-symmetric (0 ,p+ 1) tensors. How does “ d” get around the fact that the derivative of a tensor is not a tensor ? Well, if you apply the transformation law for Aµ, and the chain rule to∂ ∂xµto find the transformation law for Fµν=∂µAν−∂νAµ, you will see why: all the derivatives of the∂zν ∂xµcancel, and Fµνis abona-fide tensor of type (0 ,2). This sort of cancellation is why skew-symmetric objects are usef ul, and symmetric ones less so. Exercise 2.4 : Use axiom ii) to compute d(d(a∧b)) and confirm that it is zero. Closed and exact forms The Poincar´ e lemma. d2= 0, leads to some important terminology: i) Ap-formωis said to be closed ifdω= 0. ii) Ap-formωis said to exact ifω=dηfor some (p−1)-formη. An exact form is necessarily closed, but a closed form is not n ecessarily exact. The question of when closed ⇒exact is one involving the global topology of the space in which the forms are defined, and will be subject of chapter 4. Cartan’s formulæ It is sometimes useful to have expressions for the action of dcoupled with the evaluation of the subsequent ( p+ 1) forms. Iff,η,ω , are 0,1,2-forms, respectively, then df,dη,dω , are 1,2,3-forms. When we plug in the appropriate number of vector fields X,Y,Z , then, after some labour, we will find df(X) =Xf. (2.53) 52 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS dη(X,Y) =Xη(Y)−Yη(X)−η([X,Y]). (2.54) dω(X,Y,Z ) =Xω(Y,Z) +Yω(Z,X) +Zω(X,Y) −ω([X,Y],Z)−ω([Y,Z],X)−ω([Z,X],Y).(2.55) These formulæ, and their higher- panalogues, express din terms of geometric objects, and so make it clear that the exterior derivative is itself a geometric object, independent of any particular co-ordinate choice. Let us demonstate the correctness of the second formula. Wit hη=ηµdxµ, the left-hand side, dη(X,Y), is equal to ∂µηνdxµdxν(X,Y) =∂µην(XµYν−XνYµ). (2.56) The right hand side is equal to Xµ∂µ(ηνYν)−Yµ∂µ(ηνXν)−ην(Xµ∂µYν−Yµ∂µXν). (2.57) On using the product rule for the derivatives in the first two t erms, we find that all derivatives of the components of XandYcancel, and are left with exactly those terms appearing on left. Exercise 2.5 : Letωi,i= 1,...,r be a linearly independent set of one-forms defining a Pfaffian system (see sec. 2.2.1) in ddimensions. i) Use Cartan’s formulæ to show that the corresponding ( d−r)-dimensional distribution is involutive if and only if there is an r-by-rmatrix of 1-forms θijsuch that dωi=r/summationdisplay j=1θij∧ωj. ii) Show that the conditions in part i) are satisfied if there a rerfunctions giand an invertible r-by-rmatrix of functions fi jsuch that ωi=r/summationdisplay j=1fi jdgi. In this case foliation surfaces are given by the conditions gi(x) = const., i= 1,...,r . It is also possible, but considerably harder, to show that i) ⇒ii). Doing so would constitute a proof of Frobenius’ theorem. Exercise 2.6 : Letωbe a closed two-form, and let Null( ω) be the space of vector fields Xsuch thatω(X,) = 0. Use the Cartan formulæ to show that ifX,Y∈Null(ω), then [X,Y]∈Null(ω). 2.3. EXTERIOR CALCULUS 53 Lie Derivative of Forms Given ap-formωand a vector field X, we can form a ( p−1)-form called iXωby writing iXω(....../bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright p−1slots) =ω(pslots/bracehtipdownleft/bracehtipupright/bracehtipupleft/bracehtipdownright X,....../bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright p−1slots). (2.58) Acting on a 0-form, iXis defined to be 0. This procedure is called the interior multiplication byX. It is simply a contraction ωjij2...jp→ωkj2...jpXk, (2.59) but it is convenient to have a special symbol for this operati on. It is perhaps surprising that iXturns out to be an anti-derivation, just as is d. Ifηandω arepandqforms respectively, then iX(η∧ω) = (iXη)∧ω+ (−1)pη∧(iXω), (2.60) even though iXinvolves no differentiation. For example, if X=Xµ∂µ, then iX(dxµ∧dxν) =dxµ∧dxν(Xα∂α,), =Xµdxν−dxµXν, = (iXdxµ)∧(dxν)−dxµ∧(iXdxν). (2.61) One reason for introducing iXis that there is a nice (and profound) formula for the Lie derivative of a p-form in terms of iX. The formula is called the infinitesimal homotopy relation . It reads LXω= (diX+iXd)ω. (2.62) This formula is proved by verifying that it is true for functi ons and one- forms, and then showing that it is a derivation – in other word s that it satisfies Leibniz’ rule. From the derivation property of the Lie derivative, we immediately deduce that that the formula works for any p-form. That the formula is true for functions should be obvious: Sin ceiXf= 0 by definition, we have (diX+iXd)f=iXdf=df(X) =Xf=LXf. (2.63) 54 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS To show that the formula works for one forms, we evaluate (diX+iXd)(fνdxν) =d(fνXν) +iX(∂µfνdxµdxν) =∂µ(fνXν)dxµ+∂µfν(Xµdxν−Xνdxµ) = (Xν∂νfµ+fν∂µXν)dxµ. (2.64) In going from the second to the third line, we have interchang ed the dummy labelsµ↔νin the term containing dxν. We recognize that the 1-form in the last line is indeed LXf. To show that diX+iXdis a derivation we must apply diX+iXdtoa∧b and use the anti-derivation property of ixandd. This is straightforward once we recall that dtakes ap-form to a ( p+ 1)-form while iXtakes ap-form to a (p−1)-form. Exercise 2.7 : Let ω=1 p!ωi1...ipdxi1···dxip. Use the anti-derivation property of iXto show that iXω=1 (p−1)!ωαi2...ipXαdxi2···dxip, and so verify the equivalence of (2.58) and (2.59). Exercise 2.8 : Use the infinitesimal homotopy relation to show that Landd commute, i.e.forωap-form, we have d(LXω) =LX(dω). 2.4 Physical Applications 2.4.1 Maxwell’s Equations In relativistic3four-dimensional tensor notation the two source-free Maxw ell’s equations curlE=−∂B ∂t, divB= 0, (2.65) 3In this section we will use units in which c=/epsilon10=µ0= 1. We take the Minkowski metric to be gµν= diag (−1,1,1,1) wherex0=t,x1=x,etc. 2.4. PHYSICAL APPLICATIONS 55 reduce to the single equation ∂Fµν ∂xλ+∂Fνλ ∂xµ+∂Fλµ ∂xν= 0. (2.66) where Fµν= 0−Ex−Ey−Ez Ex 0Bz−By Ey−Bz 0Bx EzBy−Bx 0 . (2.67) The “F” is traditional, for Michael Faraday. In form language, the relativistic equation becomes the even more compact expression dF= 0, where F≡1 2Fµνdxµdxν =Bxdydz+Bydzdx+Bzdxdy+Exdxdt+Eydydt+Ezdzdt, (2.68) is a Minkowski-space 2-form. Exercise 2.9 : Verify that the source-free Maxwell equations are indeed e quiv- alent todF= 0. The equation dF= 0 is automatically satisfied if we introduce a 4-vector 1-form potential A=−φdt+Axdx+Aydy+Azdzand setF=dA. The two Maxwell equations with sources divD=ρ, curlH=j+∂D ∂t, (2.69) reduce in 4-tensor notation to the single equation ∂µFµν=Jν. (2.70) HereJµ= (ρ,j) is the current 4-vector. This source equation takes a little more work to express in fo rm language, but it can be done. We need a new concept: the Hodge “star” dual of a form. Inddimensions the “ ⋆” map takes a p-form to a ( d−p)-form. It depends on both the metric and the orientation . The latter means a canonical choice of the order in which to write our basis forms, with orderings that differ 56 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS by an even permutation being counted as the same. The full d-dimensional definition involves the Levi-Civita duality operation of ch apter 1 , combined with the use of the metric tensor to raise indices. Recall tha t√g=/radicalbig detgµν. (In Minkowski-signature metrics we should replace√gby√−g.) We define “⋆” to be a linear map ⋆:p/logicalanddisplay (T∗M)→(d−p)/logicalanddisplay (T∗M) (2.71) such that ⋆dxi1...dxipdef=1 (d−p)!√ggi1j1...gipjp/epsilon1j1···jpjp+1···jddxjp+1...dxjd.(2.72) Although this definition looks a trifle involved, computatio ns involving it are not so intimidating. The trick is to work, whenever possible , with oriented orthonormal frames. If we are in euclidean space and {e∗i1,e∗i2,...,e∗id}is an ordering of the orthonormal basis for ( T∗M)xwhose orientation is equivalent to{e∗1,e∗2,...,e∗d}then ⋆(e∗i1∧e∗i2∧···∧ e∗ip) =e∗ip+1∧e∗ip+2∧···∧ e∗id. (2.73) For example, in three dimensions, and with x,y,z, our usual Cartesian co- ordinates, we have ⋆dx =dydz, ⋆dy =dzdx, ⋆dz =dxdy. (2.74) An analogous method works for Minkowski-signature ( −,+,+,+) metrics, except that now we must include a minus sign for each negative ly normed dtfactor in the form being “starred.” Taking {dt,dx,dy,dz}as our oriented basis, we therefore find4 ⋆dxdy =−dzdt, ⋆dydz =−dxdt, ⋆dzdx =−dydt, ⋆dxdt =dydz, ⋆dydt =dzdx, ⋆dzdt =dxdy. (2.75) 4See for example: Misner, Thorn and Wheeler, Gravitation , (MTW) page 108. 2.4. PHYSICAL APPLICATIONS 57 For example, the first of these equations is derived by observ ing that (dxdy)(−dzdt) = dtdxdydz , and that there is no “ dt” in the product dxdy. The fourth fol- lows from observing that that ( dxdt)(−dydx) =dtdxdydz , but there is a negative-normed “ dt” in the product dxdt. The⋆map is constructed so that if α=1 p!αi1i2...ipdxi1dxi2···dxip, (2.76) and β=1 p!βi1i2...ipdxi1dxi2···dxip, (2.77) then α∧(⋆β) =β∧(⋆α) =/angbracketleftα,β/angbracketrightσ, (2.78) where the inner product /angbracketleftα,β/angbracketrightis defined to be the invariant /angbracketleftα,β/angbracketright=1 p!gi1j1gi2j2···gipjpαi1i2...ipβj1j2...jp, (2.79) andσis the volume form σ=√gdx1dx2···dxd. (2.80) In future we will write α⋆β forα∧(⋆β). Bear in mind that the “ ⋆” in this expression is acting βand is not some new kind of binary operation. We now apply these ideas to Maxwell. From the field-strength 2 -form F=Bxdydz+Bydzdx+Bzdxdy+Exdxdt+Eydydt+Ezdzdt, (2.81) we get a dual 2-form ⋆F=−Bxdxdt−Bydydt−Bzdzdt+Exdydz+Eydzdx+Ezdxdy. (2.82) We can check that we have correctly computed the Hodge star of Fby taking the wedge product, for which we find F ⋆F =1 2(FµνFµν)σ= (B2 x+B2 y+B2 z−E2 x−E2 y−E2 z)dtdxdydz. (2.83) Observe that the expression B2−E2is a Lorentz scalar. Similarly, from the current 1-form J≡Jµdxµ=−ρdt+jxdx+jydy+jzdz, (2.84) 58 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS we derive the dual current 3-form ⋆J=ρdxdydz−jxdtdydz−jydtdzdx−jzdtdxdy, (2.85) and check that J⋆J = (JµJµ)σ= (−ρ2+j2 x+j2 y+j2 z)dtdxdydz. (2.86) Observe that d⋆J=/parenleftbigg∂ρ ∂t+ divj/parenrightbigg dtdxdydz = 0, (2.87) expresses the charge conservation law. Writing out the terms explicitly shows that the source-cont aining Maxwell equations reduce to d⋆F=⋆J.All four Maxwell equations are therefore very compactly expressed as dF= 0, d⋆F =⋆J. Observe that current conservation d⋆J= 0 follows from the second Maxwell equation as a consequence of d2= 0. Exercise 2.10 : Show that for a p-formωindeuclidean dimensions we have ⋆⋆ω= (−1)p(d−p)ω. Show, further, that for a Minkowski metric an additional min us sign has to be inserted. (For example, ⋆⋆F =−F, even though (−1)2(4−2)= +1.) 2.4.2 Hamilton’s Equations Hamiltonian dynamics takes place in phase space , a manifold with co-ordinates (q1,...,qn,p1,...,pn). Since momentum is a naturally covariant vector5, phase space is usually the co-tangent bundle T∗Mof the configuration man- ifoldM. We are writing the indices on the p’s upstairs though, because we are considering them as co-ordinates in T∗M. We expect that you are familiar with Hamilton’s equation in t heirq,p setting. Here, we shall describe them as they appear in a mode rn book on Mechanics, such as Abrahams and Marsden’s Foundations of Mechanics , or V. I. Arnold’s Mathematical Methods of Classical Mechanics . 5To convince yourself of this, remember that in quantum mecha nics ˆpµ=−i/planckover2pi1∂ ∂xµ, and the gradient of a function is a covector. 2.4. PHYSICAL APPLICATIONS 59 Phase space is an example of a symplectic manifold , a manifold equiped with a symplectic form — a non-degenerate 2-form field ω=1 2ωijdxidxj. (2.88) Recall that the word closed means that dω= 0.Non-degenerate means that for any point xthe statement that ω(X,Y) = 0 for all vectors Y∈TMx implies that X= 0 at that point (or equivalently that for all xthe matrix ωij(x) has an inverse ωij(x)). Given a Hamiltonian functionHon our symplectic manifold, we define a velocity vector-field vHby solving dH=−ivHω=−ω(vH,) (2.89) forvH. If the symplectic form is ω=dp1dq1+dp2dq2+···+dpndqn, this is nothing but a fancy form of Hamilton’s equations. To see this , we write dH=∂H ∂qidqi+∂H ∂pidpi(2.90) and use the customary notation ( ˙ qi,˙pi) for the velocity-in-phase-space com- ponents, so that vH= ˙qi∂ ∂qi+ ˙pi∂ ∂pi. (2.91) Now we work out ivHω=dpidqi( ˙qj∂qj+ ˙pj∂pj,) = ˙pidqi−˙qidpi, (2.92) so, comparing coefficients of dpianddqion the two sides of dH=−ivHω, we read off ˙qi=∂H ∂pi,˙pi=−∂H ∂qi. (2.93) Darboux’ theorem , which we will not try to prove, says that for any point x we can always find co-ordinates p,q, valid in some neigbourhood of x, such thatω=dp1dq1+dp2dq2+···dpndqn. Given this fact, it is not unreasonable to think that there is little to gained by using the abstract d ifferential-form language. In simple cases this is so, and the traditional met hods are fine. 60 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS It may be, however, that the neigbourhood of xwhere the Darboux co- ordinates work is not the entire phase space, and we need to co ver the space with overlapping p,qco-ordinate charts. Then, what is a pin one chart will usually be a combination of p’s andq’s in another. In this case, the traditional form of Hamilton’s equations loses its appeal i n comparison to the co-ordinate-free dH=−ivHω. Given two functions H1,H2we can define their Poisson bracket{H1,H2}. Its importance lies in Dirac’s observation that the passage from classical mechanics to quantum mechanics is accomplished by replacin g the Poisson bracket of two quantities, AandB, with the commutator of the correspond- ing operators ˆA, and ˆB: i[ˆA,ˆB]←→ /planckover2pi1{A,B}+O/parenleftbig /planckover2pi12/parenrightbig . (2.94) We define the Poisson bracket by6 {H1,H2}def=dH2 dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle H1=vH1H2. (2.95) Now,vH1H2=dH2(vH1), and Hamilton’s equations say that dH2(vH1) = ω(vH1,vH2). Thus, {H1,H2}=ω(vH1,vH2). (2.96) The skew symmetry of ω(vH1,vH2) shows that despite the asymmetrical ap- pearance of the definition we have skew symmetry: {H1,H2}=−{H2,H1}. Moreover, since vH1(H2H3) = (vH1H2)H3+H2(vH1H3), (2.97) the Poisson bracket is a derivation: {H1,H2H3}={H1,H2}H3+H2{H1,H3}. (2.98) Neither the skew symmetry nor the derivation property requi re the con- dition that dω= 0. What does need ωto be closed is the Jacobi identity : {{H1,H2},H3}+{{H2,H3},H1}+{{H3,H1},H2}= 0. (2.99) 6Our definition differs in sign from the traditional one, but ha s the advantage of mini- mizing the number of minus signs in subsequent equations. 2.4. PHYSICAL APPLICATIONS 61 We establish Jacobi by using Cartan’s formula in the form dω(vH1,vH2,vH3) =vH1ω(vH2,vH3) +vH2ω(vH3,vH1) +vH3ω(vH1,vH2) −ω([vH1,vH2],vH3)−ω([vH2,vH3],vH1)−ω([vH3,vH1],vH2). (2.100) It is relatively straight-forward to interpret each term in the first of these lines as Poisson brackets. For example, vH1ω(vH2,vH3) =vH1{H2,H3}={H1,{H2,H3}}. (2.101) Relating the terms in the second line to Poisson brackets req uires a little more effort. We proceed as follows: ω([vH1,vH2],vH3) =−ω(vH3,[vH1,vH2]) =dH3([vH1,vH2]) = [vH1,vH2]H3 =vH1(vH2H3)−vH2(vH1H3) ={H1,{H2,H3}}−{H2,{H1,H3}} ={H1,{H2,H3}}+{H2,{H3,H1}}.(2.102) Adding everything togther now shows that 0 =dω(vH1,vH2,vH3) =−{{H1,H2},H3}−{{H2,H3},H1}−{{H3,H1},H2}.(2.103) If we rearrange the Jacobi identity as {H1,{H2,H3}}−{H2,{H1,H3}}={{H1,H2},H3}, (2.104) we see that it is equivalent to [vH1,vH2] =v{H1,H2}. The algebra of Poisson brackets is therefore homomorphic to the algebra of the Lie brackets. The correspondence is not an isomorphism , however: the assignment H/mapsto→vHfails to be one-to-one because constant functions map to the zero vector field. Exercise 2.11 : Use the infinitesimal homotopy relation, to show that LvHω= 0, wherevHis the vector field corresponding to H. Suppose now that the phase space is 2ndimensional. Show that in local Darboux co-ordinates the 2 n-form ωn/n! is, up to a sign, the phase-space volume element dnpdnq. Show that LvHωn/n! = 0 and that this result is Liouville’s theorem on the conservation of phase-space volume. 62 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS The classical mechanics of spin It is sometimes said in books on quantum mechanics that the sp in of an elec- tron, or other elementary particle, is a purely quantum conc ept and cannot be described by classical mechanics. This statement is fals e, but spin isthe simplest system in which traditional physicist’s methods b ecome ugly and it helps to use the modern symplectic language. A “spin” Scan be regarded as a fixed length vector that can point in any direction in R3. We will take it to be of unit length so that its components are Sx= sinθcosφ, Sy= sinθsinφ, Sz= cosθ, (2.105) whereθandφare polar co-ordinates on the two-sphere S2. The surface of the sphere turns out to be both the configuratio n space and the phase space. In particular the phase space for a spin i snotthe cotangent bundle of the configuration space. This has to be so : we learned from Niels Bohr that a 2 n-dimensional phase space contains roughly one quantum state for every /planckover2pi1nof phase-space volume. A cotangent bundle always has infinite volume, so its corresponding Hilbert spa ce is necessarily infinite dimensional. A quantum spin, however, has a finite-dimensional Hilbert space so its classical phase space must have a finite t otal volume. This finite-volume phase space seems un-natural in the tradi tional view of mechanics, but it fits comfortably into the modern symplecti c picture. We want to treat all points on the sphere alike, and so it is nat ural to take the symplectic 2-form to be proportional to the element of ar ea. Suppose that ω= sinθdθdφ . We could write ω=dcosθdφand regard φas “q” and cosθ as “p’ (Darboux’ theorem in action!), but this identification is s ingular at the north and south poles of the sphere, and, besides, it obscure s the spherical symmetry of the problem, which is manifest when we think of ωasd(area). Let us take our hamiltonian to be H=BSx, corresponding to an applied magnetic field in the xdirection, and see what Hamilton’s equations give for the motion. First we take the exterior derivative d(BSx) =B(cosθcosφdθ−sinθsinφdφ). (2.106) This is to be set equal to −ω(vBSx,) =vθ(−sinθ)dφ+vφsinθdθ. (2.107) 2.5. COVARIANT DERIVATIVES 63 Comparing coefficients of dθanddφ, we get v(BSx)=vθ∂θ+vφ∂φ=B(sinφ∂θ+ cosφcotθ∂φ), (2.108) i.e.Btimes the velocity vector for a rotation about the xaxis. This velocity field therefore describes a steady Larmor precession of the s pin about the applied field. This is exactly the motion predicted by quantu m mechanics. Similarly, setting B= 1, we find vSy=−cosφ∂θ+ sinφcotθ∂φ, vSz=−∂φ. (2.109) From these velocity fields we can compute the Poisson bracket s: {Sx,Sy}=ω(vSx,vSy) = sinθdθdφ (sinφ∂θ+ cosφcotθ∂φ,−cosφ∂θ+ sinφcotθ∂φ) = sinθ(sin2φcotθ+ cos2φcotθ) = cosθ=Sz. Repeating the exercise leads to {Sx,Sy}=Sz, {Sy,Sz}=Sx, {Sz,Sx}=Sy. (2.110) These Poisson brackets for our classical “spin” are to be com pared with the commutation relations [ ˆSx,ˆSy] =i/planckover2pi1ˆSzetc.for the quantum spin operators ˆSi. 2.5 Covariant Derivatives Covariant derivatives are a general class of derivatives th at act on sections of a vector or tensor bundle over a manifold. We will begin by c onsidering derivatives on the tangent bundle, and in the exercises indi cate how the idea generalizes to other bundles. 64 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS 2.5.1 Connections The Lie and exterior derivatives require no structure beyon d that which comes for free with our manifold. Another type of derivative that can act on tangent-space vectors and tensors is the covariant derivative ∇X≡Xµ∇µ. This requires an additional mathematical object called an affine connection . The covariant derivative is defined by: i) Its action on scalar functions as ∇Xf=Xf. (2.111) ii) Its action a basis set of tangent-vector fields ea(x) =eµ a(x)∂µ(a local frame, or vielbein7) by introducing a set of functions ωi jk(x) and setting ∇ekej=eiωi jk. (2.112) ii) Extending this definition to any other type of tensor by re quiring∇X to be a derivation. iii) Requiring that the result of applying ∇Xto a tensor is a tensor of the same type. The set of functions ωi jk(x) is the connection . In any local co-ordinate chart we can choose them at will, and different choices define differe nt covariant derivatives. (There may be global compatibility constrain ts, however, which appear when we assemble the charts into an atlas.) Warning : Despite having the appearance of one, ωi jkisnota tensor. It transforms inhomogeneously under a change of frame or co-or dinates — see equation (2.131). We can, of course, take as our basis vectors the co-ordinate v ectors eµ≡ ∂µ. When we do this it is traditional to use the symbol Γ for the co -ordinate frame connection instead of ω. Thus, ∇µeν≡∇eµeν=eλΓλ νµ. (2.113) The numbers Γλνµare often called Christoffel symbols . As an example consider the covariant derivative of a vector fνeν. Using the derivation property we have ∇µ(fνeν) = (∂µfν)eν+fν∇µeν = (∂µfν)eν+fνeλΓλ νµ =eν/braceleftbig ∂µfν+fλΓν λµ/bracerightbig . (2.114) 7In practice viel, “many”, is replaced by the appropriate German numeral: ein-, zwei-, drei-, vier-, f¨ unf-, ..., indicating the dimension. The word beinmeans “leg.” 2.5. COVARIANT DERIVATIVES 65 In the first line we have used the defining property that ∇eµacts on the functionsfνas∂µ, and in the last line we interchanged the dummy indices νandλ. We often abuse the notation by writing only the components, and set ∇µfν=∂µfν+fλΓν λµ. (2.115) Similarly, acting on the components of a mixed tensor, we wou ld write ∇µAα βγ=∂µAα βγ+ Γα λµAλ βγ−Γλ βµAα λγ−Γλ γµAα βλ. (2.116) When we use this notation, we are no longer regarding the tens or components as “functions.” Observe that the plus and minus signs in (2.116) are required so that, for example, the covariant derivative of the scalar function fαgαis ∇µ(fαgα) =∂µ(fαgα) = (∂µfα)gα+fα(∂µgα) =/parenleftbig ∂µfα−fλΓλ αµ/parenrightbig gα+fα/parenleftbig ∂µgα+gλΓα λµ/parenrightbig = (∇µfα)gα+fα(∇µgα), (2.117) and so satisfies the derivation property. Parallel transport We have defined the covariant derivative viaits formal calculus properties. It has, however, a geometrical interpretation. As with the L ie derivative, in order to compute the derivative along Xof the vector field Y, we have to somehow carry the vector Y(x) from the tangent space TMxto the tangent spaceTMx+/epsilon1X, where we can subtract it from Y(x+/epsilon1X) . The Lie derivative carriesYalong with the Xflow. The covariant derivative implicitly carries Yby “parallel transport”. If γ:s/mapsto→xµ(s) is a parameterized curve with tangent vector Xµ∂µ, where Xµ=dxµ ds, (2.118) then we say that the vector field Y(xµ(s)) isparallel transported along the curveγif ∇XY= 0, (2.119) 66 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS at each point xµ(s). Thus, a vector that in the vielbein frame eiatxhas components Yiwill, after being parallel transported to x+/epsilon1X, end up com- ponents Yi−/epsilon1ωi jkYjXk. (2.120) In a co-ordinate frame, after parallel transport through an infinitesimal dis- placementδxµ, the vector Yν∂νwill have components Yν→Yν−Γν λµYλδxµ, (2.121) and so δxµ∇µYν=Yν(xµ+δxµ)−{Yν(x)−Γν λµYλδxµ} =δxµ{∂µYν+ Γν λµYλ}. (2.122) Curvature and Torsion As we said earlier, the connection ωi jk(x) is not itself a tensor. Two important quantities which aretensors, are associated with ∇X: i) The torsion T(X,Y) =∇XY−∇YX−[X,Y]. (2.123) The quantity T(X,Y) is a vector depending linearly on X,Y, soTat the pointxis a mapTMx×TMx→TMx, and so a tensor of type (1,2). In a co-ordinate frame it has components Tλ µν= Γλ µν−Γλ νµ. (2.124) ii) The Riemann curvature tensor R(X,Y)Z=∇X∇YZ−∇Y∇ZZ−∇ [X,Y]Z. (2.125) The quantity R(X,Y)Zis also a vector, so R(X,Y) is a linear map TMx→TMx, and thusRitself is a tensor of type (1,3). Written out in a co-ordinate frame, we have Rα βµν=∂µΓα βν−∂νΓα βµ+ Γα λµΓλ βν−Γα λνΓλ βµ. (2.126) If our manifold comes equipped with a metric tensor gµν(and is thus aRiemann manifold ), and if we require both that T= 0 and∇µgαβ= 0, 2.5. COVARIANT DERIVATIVES 67 then the connection is uniquely determined, and is called th eRiemann , or Levi-Civita , connection. In a co-ordinate frame it is given by Γα µν=1 2gαλ(∂µgλν+∂νgµλ−∂λgµν). (2.127) This is the connection that appears in General Relativity. The curvature tensor measures the degree of path dependence in parallel transport: if Yν(x) is parallel transported along a path γ:s/mapsto→xµ(s) from atob, and if we deform γso thatxµ(s)→xµ(s) +δxµ(s) while keeping the endpointsa,bfixed, then δYα(b) =−/integraldisplayb aRα βµν(x)Yβ(x)δxµdxν. (2.128) IfRαβµν≡0 then the effect of parallel transport from atobwill be indepen- dent of the route taken. The geometric interpretation of Tµνis less transparent. On a two-dimensional surface a connection is torsion free when the tangent space “ rolls without slipping” along the curve γ. Exercise 2.12 :Metric compatibility . Show that the Riemann connection Γαµν=1 2gαλ(∂µgλν+∂νgµλ−∂λgµν). follows from the torsion-free condition Γαµν= Γανµtogether with the metric compatibility condition ∇µgαβ≡∂µgαβ−Γναµgνβ−Γναµgαν= 0. Show that “metric compatibility” means that that the operat ion of raising or lowering indices commutes with covariant derivation. Exercise 2.13 :Geodesic equation . Letγ:s/mapsto→xµ(s) be a parametrized path fromatob. Show that the Euler-Lagrange equation that follows from minimizing the distance functional S(γ) =/integraldisplayb a/radicalbig gµν˙xµ˙xνds, where the dots denote differentiation with respect to the par ameters, is d2xµ ds2+ Γµαβdxα dsdxβ ds= 0. Here Γµαβis the Riemann connection (2.127). 68 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS Exercise 2.14 : Show that if Aµis a vector field then, for the Riemann connec- tion, ∇µAµ=1√g∂√gAµ ∂xµ. In other words, show that Γααµ=1√g∂√g ∂xµ. Deduce that the Laplacian acting on a scalar field φcan be defined by setting either ∇2φ=gµν∇µ∇νφ, or ∇2φ=1√g∂ ∂xµ/parenleftbigg√ggµν∂φ ∂xν/parenrightbigg , the two definitions being equivalent. 2.5.2 Cartan’s Form Viewpoint Lete∗j(x) =e∗j µ(x)dxµbe the basis of one-forms dual to the vielbein frame ei(x) =eµ i(x)∂µ. Since δi j=e∗i(ej) =e∗j µeµ i, (2.129) the matrices e∗j µandeµ iare inverses of one-another. We can use them to change from roman vielbein indices to greek co-ordinate fra me indices. For example: gij=g(ei,ej) =eµ igµνeν j, (2.130) and ωi jk=e∗i ν(∂µeν j)eµ k+e∗i λeν jeµ kΓλ νµ. (2.131) Cartan regards the connection as being a matrix Ωof one-forms with entriesωi j=ωi jµdxµ. In this language equation (2.112) becomes ∇Xej=eiωi j(X). (2.132) Cartan’s viewpoint separates off the index µ, which refers to the direction δxµ∝Xµin which we are differentiating, from the matrix indices iand jthat act on the components of the vector or tensor being differ entiated. This separation becomes very natural when the vector space s panned by the 2.5. COVARIANT DERIVATIVES 69 ei(x) is no longer the tangent space, but some other “internal” ve ctor space attached to the point x. Such internal spaces are common in physics, an im- portant example being the “colour space” of gauge field theor ies. Physicists, following Hermann Weyl, call a connection on an internal spa ce a “gauge po- tential.” To mathematicians it is simply a connection on the vector bundle that has the internal spaces as its fibres. Cartan also regards the torsion Tand curvature Ras forms; in this case vector- and matrix-valued two-forms, respectively, with e ntries Ti=1 2Ti µνdxµdxν, (2.133) Ri k=1 2Ri kµνdxµdxν. (2.134) In his form language the equations defining the torsion and cu rvature become Cartan’s structure equations : de∗i+ωi j∧e∗j=Ti, (2.135) and dωi k+ωi j∧ωj k=Ri k. (2.136) The last equation can be written more compactly as dΩ+Ω∧Ω=R. (2.137) From this, by taking the exterior derivative, we obtain the Bianchi identity dR−R∧Ω+Ω∧R= 0. (2.138) On a Riemann manifold, we can take the vielbein frame eito be orthonor- mal. In this case the roman-index metric gij=g(ei,ej) becomesδij. There is then no distinction between covariant and contravariant roman indices, and the connection and curvature forms, Ω,R, being infinitesimal rotations, become skew symmetric matrices: ωij=−ωji, Rij=−Rji. (2.139) 70 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS 2.6 Further Exercises and Problems Exercise 2.15 : Consider the vector fields X=y∂x,Y=∂yinR2. Find the flows associated with these fields, and use them to verify the s tatements made in section 2.2.1 about the geometric interpretation of the L ie bracket. Exercise 2.16 : Show that the pair of vector fields Lz=x∂y−y∂xandLy= z∂x−x∂zinR3is in involution wherever they are both non-zero. Show furth er that the general solution of the system of partial differenti al equations (x∂y−y∂x)f= 0, (x∂z−z∂x)f= 0, inR3isf(x,y,z) =F(x2+y2+z2), whereFis an arbitrary function. Exercise 2.17 : In the rolling conditions (2.26) we are using the “ Y” convention for Euler angles. In this convention θandφare the usual spherical polar co- ordinate angles with respect to the space-fixed xyzaxes. They specify the direction of the body-fixed Zaxis about which we make the final ψrotation. θ φz yxZ Y YXψ Figure 2.7: Euler angles: we first rotate the ball through an angle φabout thezaxis, thus taking y→Y/prime, then through θaboutY/prime, and finally through ψaboutZ, so taking Y/prime→Y. a) Show that (2.26) are indeed the no-slip rolling condition s ˙x=ωy, ˙y=−ωx, 0 =ωz, 2.6. FURTHER EXERCISES AND PROBLEMS 71 where (ωx,ωy,ωz) are the components of the ball’s angular velocity in thexyzspace-fixed frame. b) Solve the three constraints in (2.26) so as to obtain the ve ctor fields (2.27). c) Show that [rollx,rolly] =−spinz, wherespinz≡∂φ, corresponds to a rotation about a vertical axis through the point of contact. This is a new motion, being forbidden by theωz= 0 condition. d) Show that [spinz,rollx] = spinx, [spinz,rolly] = spiny, where the new vector fields spinx≡ −(rolly−∂y), spiny≡(rollx−∂x), correspond to rotations of the ball about the space-fixed xandyaxes through its centre, and with the centre of mass held fixed. We have generated five independent vector fields from the orig inal two. There- fore, by sufficient rolling to-and-fro, we can position the ba ll anywhere on the table, and in any orientation. Exercise 2.18 : The semi-classical dynamics of charge −eelectrons in a mag- netic solid are governed by the equations8 ˙r=∂/epsilon1(k) ∂k−˙k×Ω, ˙k=−∂V ∂r−e˙r×B. Herekis the Bloch momentum of the electron, ris its position, /epsilon1(k) its band energy (in the extended-zone scheme), and B(r) is the external magnetic field. The components Ω iof the Berry curvature Ω(k) are given in terms of the periodic part|u(k)/angbracketrightof the Bloch wavefunctions of the band by Ωi=i/epsilon1ijk1 2/parenleftBigg/angbracketleftBigg ∂u ∂kj/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle∂u ∂kk/angbracketrightBigg −/angbracketleftBigg ∂u ∂kk/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle∂u ∂kj/angbracketrightBigg/parenrightBigg . 8M. C. Chang, Q. Niu, Phys. Rev. Lett. 75(1995) 1348. 72 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS The only property of Ω(k) needed for the present problem, however, is that divkΩ= 0. a) Show that these equations are Hamiltonian, with H(r,k) =/epsilon1(k) +V(r) and with ω=dkidxi−e 2/epsilon1ijkBi(r)dxjdxk+1 2/epsilon1ijkΩi(k)dkjdkk. as the symplectic form.9 b) Confirm that the ωdefined in part b) is closed, and that the Poisson brackets are given by {xi,xj}=−/epsilon1ijkΩk (1 +eB·Ω), {xi,kj}=−δij+ ΩiBj (1 +eB·Ω), {ki,kj}=/epsilon1ijkBk (1 +eB·Ω). c) Show that the conserved phase-space volume ω3/3! is equal to (1 +eB·Ω)d3kd3x, instead of the na¨ ıvely expected d3kd3x. The following pair of exercises show that Cartan’s expressi on for the curva- ture tensor remains valid for covariant differentiation in “ internal” spaces. There is, however, no natural concept analogous to the torsi on tensor for internal spaces. Exercise 2.19 :Non-abelian gauge fields as matrix-valued forms . In a non- abelian Yang-Mills gauge theory, such as QCD, the vector pot ential A=Aµdxµ is matrix-valued, meaning that the components Aµare matrices which do not necessarily commute with each other. (These matrices are el ements of the Lie 9C. Duval, Z. Horv´ ath, P. A. Horv´ athy, L. Martina, P. C. Stic hel,Modern Physics Letters B 20(2006) 373. 2.6. FURTHER EXERCISES AND PROBLEMS 73 algebra of the gauge group, but we won’t need this fact here.) The matrix- valued curvature, or field-strength, 2-form Fis defined by F=dA+A2=1 2Fµνdxµdxν. Here a combined matrix and wedge product is to be understood: (A2)a b≡Aac∧Acb=AacµAcbνdxµdxν. i) Show that A2=1 2[Aµ,Aν]dxµdxν, and hence show that Fµν=∂µAν−∂νAµ+ [Aµ,Aν]. ii) Define the gauge-covariant derivatives ∇µ=∂µ+Aµ, and show that the commutator [ ∇µ,∇ν] of two of these is equal to Fµν. Show further that if X,Yare two vector fields with Lie bracket [ X,Y] and∇X≡Xµ∇µ, then F(X,Y) = [∇X,∇Y]−∇[X,Y]. iii) Show that Fobeys the Bianchi identity dF−FA+AF= 0. Again wedge and matrix products are to be understood. This eq uation is the non-abelian version of the source-free Maxwell equat iondF= 0. iv) Show that, in any number of dimensions, the Bianchi ident ity implies that the 4-form tr ( F2) is closed, i.e.thatdtr (F2) = 0. Similarly show that the 2n-form tr (Fn) is closed. (Here the “tr” means a trace over the roman matrix indices, and not over the greek space-time indi ces.) v) Show that, tr (F2) =d/braceleftbigg tr/parenleftbigg AdA+2 3A3/parenrightbigg/bracerightbigg . The 3-form tr ( AdA+2 3A3) is called a Chern-Simons form. Exercise 2.20 :Gauge transformations . Here we consider how the matrix- valued vector potential transforms when we make a change of g auge. In other words, we seek the non-abelian version of Aµ→Aµ+∂µφ. 74 CHAPTER 2. DIFFERENTIAL CALCULUS ON MANIFOLDS i) Letgbe an invertable matrix, and δga matrix describing a small change ing. Show that the corresponding change in the inverse matrix is given byδ(g−1) =−g−1(δg)g−1. ii) Show that under the gauge transformation A→Ag≡g−1Ag+g−1dg, we haveF→g−1Fg. (Hint: The labour is minimized by exploiting the covariant derivative identity in part ii) of the previous ex ercise). iii) Deduce that tr ( Fn) isgauge invariant . iv) Show that a necessary condition for the matrix-valued ga uge fieldAto be “pure gauge”, i.e.for there to be a position dependent matrix gsuch thatA=g−1dg, is thatF= 0, where Fis the curvature two-form of the previous exercise. In a gauge theory based on a Lie group G, the matrices gwill be elements of the group, or, more generally, they will form a matrix repres entation of the group. Chapter 3 Integration on Manifolds One usually thinks of integration as requiring measure – a notion of volume, and hence of size and length, and so a metric . A metric however is not required for integrating differential forms. They come pre- equipped with whatever notion of length, area, or volume is required. 3.1 Basic Notions 3.1.1 Line Integrals Consider, for example, the form df. We want to try to give a meaning to the symbol I1=/integraldisplay Γdf. (3.1) Here Γ is a path in our space starting at some point P0and ending at the point P1. Any reasonable definition of I1should end up with the answer we would immediately write down if we saw an expression like I1in an elementary calculus class. This answer is I1=/integraldisplay Γdf=f(P1)−f(P0). (3.2) No notion of a metric is needed here. There is however a geomet ric picture of what we have done. We draw in our space the surfaces ...,f(x) =−1,f(x) = 0,f(x) = 1,..., and perhaps fill in intermediate values if necessary. We then start at P0and travel from there to P1, keeping track of how many of 75 76 CHAPTER 3. INTEGRATION ON MANIFOLDS these surfaces we pass through (with sign -1, if we pass back t hrough them). The integral of dfis this number. Figure 3.1 illustrates a case in which/integraltext Γdf= 5.5−1.5 = 4. P1 f=1 2 3 4 5 6Γ P0 Figure 3.1: The integral of a one-form What we have defined is a signed integral . If we parameterise the path as x(s), 0≤s≤1, and with x(0) =P0,x(1) =P1we have I1=/integraldisplay1 0/parenleftbiggdf ds/parenrightbigg ds (3.3) where the right hand side is an ordinary one-variable integr al. It is important that we did not write/vextendsingle/vextendsingledf ds/vextendsingle/vextendsinglein this integral. The absence of the modulus sign ensures that if we partially retrace our route, so that we pas s over some part of Γ three times—twice forward and once back—we obtain the sa me answer as if we went only forward. 3.1.2 Skew-symmetry and Orientations What about integrating 2 and 3-forms? Why the skew-symmetry ? To answer these questions, think about assigning some sort of “area” i nR2to the par- allelogram defined by the two vectors x,y. This is going to be some function of the two vectors. Let us call it ω(x,y). What properties do we demand of this function? There are at least three: i) Scaling: If we double the length of one of the vectors, we ex pect the area to double. Generalizing this, we demand ω(λx,µy) = (λµ)ω(x,y). (Note that we are not putting modulus signs on the lengths, so we are allowing negative “areas”, and for the sign to change when we reverse the direction of a vector.) 3.1. BASIC NOTIONS 77 ii) Additivity: The drawing in figure 3.2 shows that we ought t o have ω(x1+x2,y) =ω(x1,y) +ω(x2,y), (3.4) similarly for the second slots. x x yx+x212 1 Figure 3.2: Additivity of ω(x,y). iii) Degeneration: If the two sides coincide, the area shoul d be zero. Thus ω(x,x) = 0. The first two properties, show that ωshould be a multilinear form. The third shows that it must be skew-symmetric! 0 =ω(x+y,x+y) =ω(x,x) +ω(x,y) +ω(y,x) +ω(y,y) =ω(x,y) +ω(y,x). (3.5) So ω(x,y) =−ω(y,x). (3.6) These are exactly the properties possessed by a 2-form. Simi larly, a 3-form outputs a volume element. These volume elements are oriented . Remember that an orientation of a set of vectors is a choice of order in which to write them. If we interchange two vectors, the orientation changes sign. We do not disting uish orientations related by an even number of interchanges. A p-form assigns a signed ( ±) p-dimensional volume element to an orientated set of vectors . If we change the orientation, we change the sign of the volume element. Orientable and Non-orientable Manifolds In the classic video game Asteroids you could select periodic boundary con- ditions so that your spaceship would leave the right-hand si de of the screen 78 CHAPTER 3. INTEGRATION ON MANIFOLDS a) b)T RP2 2 Figure 3.3: A spaceship leaves one side of the screen and returns on the ot her with a) torus boundary conditions, b) projective-plane bou ndary conditions. Observe how, in case b), the spaceship has changed from being left handed to being right-handed. and re-appear on the left. The game universe was topological ly a torusT2. Suppose that we modify the game code so that each bit of the spa ceship re-appears at the point diametrically opposite the point it left. This does not seem like a drastic change until you play a game with a left-ha nd-drive (US) spaceship. If you send the ship off the screen and watch as it re -appears on the opposite side, you will observe the ship transmogrify into a right-hand-drive (British) craft. If we ourselves made such an excursion, we w ould end up starving to death because all our left-handed digestive enz ymes would have been converted to right-handed ones. The manifold we have co nstructed is topologically equivalent to the real projective plane RP2. The lack of a global notion of being left or right-handed makes it an example of a non-orientable manifold. A manifold or surface is orientable if we can choose a global orientation for the tangent bundle. The simplest way to do this would be to find a smoothly varying set of basis-vector fields, eµ(x), on the surface and define the orientation by chosing an order, e1(x),e2(x),...,ed(x), in which to write them. In general, however, a globally-defined smooth basis w ill not exist (try to construct one for the two-sphere, S2!). We will, however, be able to find a continously varying orientated basis e(i) 1(x),e(i) 2(x),...,e(i) d(x) for each member, labelled by ( i), of an atlas of coordinate charts. We should chose 3.2. INTEGRATING P-FORMS 79 the charts so the intersection of any pair forms a connected s et. Assuming that this has been done, the orientation of pair of overlappi ng charts is said to coincide if the determinant, det A, of the map e(i) µ=Aν µe(j) νrelating the bases in the region of overlap, is positive.1If bases can be chosen so that all overlap determinants are positive, the manifold is orientable and the selected bases define the orientation. If bases cannot be so chosen, th e manifold or surface is non-orientable . Exercise 3.1 : Consider a three-dimensional ballB3with diametrically oppo- site points of its surface identified. What would happen to an aircraft flying through the surface of the ball? Would it change handedness, turn inside out, or simply turn upside down? Is this ball an orientable 3-mani fold? 3.2 Integrating p-Forms Ap-form is naturally integrated over an oriented p-dimensional surface or manifold. Rather than start with an abstract definition, We w ill first explain this pictorially, and then translate the pictures into math ematics. 3.2.1 Counting Boxes To visualize integrating 2-forms let us try to make sense of /integraldisplay Ωdfdg, (3.7) where Ω is an oriented region embedded in three dimensions. T he surfaces f=const. andg=const. break the space up into a series of tubes. The oriented surface Ω cuts these tubes in a two-dimensional mes h of (oriented) parallelograms. 1The determinant will have the same sign in the entire overlap region. If it did not, continuity and connectedness would force it to be zero somew here, implying that one of the putative bases was not linearly independent there 80 CHAPTER 3. INTEGRATION ON MANIFOLDS f=1f=2f=3g=2g=3g=4 Ω Figure 3.4: The integration region cuts the tubes into parallelograms. We define an integral by counting how many parallelograms (in cluding frac- tions of a parallelogram) there are, taking the number to be p ositive if the parallelogram given by the mesh is oriented in the same way as the surface, and negative otherwise. To compute /integraldisplay Ωhdfdg (3.8) we do the same, but weight each parallelogram, by the value of hat that point. The integral/integraltext Ωfdxdy , over a region in R2thus ends up being the number we would compute in a multivariate calculus class, bu t the integral/integraltext Ωfdydx , would be minus this. Similarly we compute /integraldisplay Ξdfdgdh (3.9) of the 3-form dfdgdh over the oriented volume Ξ, by counting how many boxes defined by the surfaces f,g,h = constant, are included in Ξ. An equivalent way of thinking of the integral of a p-form uses its definition as a skew-symmetric p-linear function. Accordingly we evaluate I2=/integraldisplay Ωω, (3.10) whereωis a 2-form, and Ω is an oriented 2-surface, by plugging vecto rs intoω. We tile the surface Ω with collection of (small) parallelog rams, each defined by an oriented pair of basis vectors v1andv2. 3.2. INTEGRATING P-FORMS 81 Ωx1v2v Figure 3.5: We tile Ωwith small oriented parallelograms and compute/summationtext x∈Ωω(v1(x),v2(x)). At each base point xwe insert these vectors into the 2-form (in the order spec- ified by their orientation) to get ω(v1,v2), and then sum the resulting num- bers to get I2. Similarly, we integrate p-form over an oriented p-dimensional region by decomposing the region into infinitesimal p-dimensional oriented parallelepipeds, inserting their defining vectors into the form, and summing their contributions. 3.2.2 Relation to conventional integrals The previous section explained how to think pictorially abo ut the integral. Here we interpret the pictures as multi-variable calculus. We begin by motivating our recipe by considering a change of v ariables in an integral in R2. Suppose we set x1=x(y1,y2),x2=x2(y1,y2) in I4=/integraldisplay Ωf(x)dx1dx2(3.11) and use dx1=∂x1 ∂y1dy1+∂x1 ∂y2dy2, dx2=∂x2 ∂y1dy1+∂x2 ∂y2dy2. (3.12) Sincedy1dy2=−dy2dy1, we have dx1dx2=/parenleftbigg∂x1 ∂y1∂x2 ∂y2−∂x2 ∂y1∂x1 ∂y2/parenrightbigg dy1dy2. (3.13) 82 CHAPTER 3. INTEGRATION ON MANIFOLDS Thus /integraldisplay Ωf(x)dx1dx2=/integraldisplay Ω/primef(x(y))∂(x1,x2) ∂(y1,y2)dy1dy2(3.14) where∂(x1,y1) ∂(y1,y2)is the Jacobean determinant ∂(x1,y1) ∂(y1,y2)≡/parenleftbigg∂x1 ∂y1∂x2 ∂y2−∂x2 ∂y1∂x1 ∂y2/parenrightbigg , (3.15) and Ω/primethe integration region in the new variables. There is theref ore no need to include an explicit Jacobean factor when changing variab les in an integral of ap-form over a p-dimensional space—it comes for free with the form. This observation leads us to the general prescription: To ev aluate/integraltext Ωω, the integral of a p-form ω=1 p!ωµ1µ2...µpdxµ1···dxµp(3.16) over the region Ω of a pdimensional surface in a d≥pdimensional space, substitute a paramaterization x1=x1(ξ1,ξ2,...,ξp), ... xd=xd(ξ1,ξ2,...,ξp), (3.17) of the surface into ω. Next, use dxµ=∂xµ ∂ξidξi, (3.18) so that ω→ω(x(ξ))i1i2...ip∂xi1 ∂ξ1···∂xip ∂ξpdξ1···dξp, (3.19) which we regard as a p-form on Ω. (Our customary 1 /p! is absent here because we have chosen a particular order for the dξ’s.) Then /integraldisplay Ωωdef=/integraldisplay Ωω(x(ξ))i1i2...ip∂xi1 ∂ξ1···∂xip ∂ξpdξ1···dξp, (3.20) where the right hand side is an ordinary multiple integral. T his recipe is a generalization of the formula (3.3) which reduced the integ ral of a one-form 3.2. INTEGRATING P-FORMS 83 to an ordinary single-variable integral. Because the appro priate Jacobean factor appears automatically, the numerical value of the in tegral does not depend on the choice of parameterization of the surface. Example : To integrate the 2-form xdydz over the surface of a two dimen- sional sphere of radius R, we parameterize the surface with polar angles as x=Rsinφsinθ, y=Rcosφsinθ, z=Rcosθ. (3.21) Then dy=−Rsinφsinθdφ+Rcosφcosθdθ, dz=−Rsinθdθ, (3.22) and so xdydz =R3sin2φsin3θdφdθ. (3.23) We therefore evaluate /integraldisplay spherexdydz =R3/integraldisplay2π 0/integraldisplayπ 0sin2φsin3θdφdθ =R3/integraldisplay2π 0sin2φdφ/integraldisplayπ 0sin3θdθ =R3π/integraldisplay1 −1(1−cos2θ)dcosθ =4 3πR3. (3.24) The volume form Although we do not need any notion of length to integrate a diff erential form, ap-dimensional surface embedded or immersed in Rddoes inherit a distance scale from the RdEuclidean metric, and this is used to define the area or volume of the surface. When the Cartesian co-ordinat esx1,...,xd of a point in the surface are given as xa(ξ1,...,ξp), where the ξ1,...,ξp,are co-ordinates on the surface, then the inherited, or induced , metric is “ds2”≡g(,)≡gµνdξµ⊗dξν(3.25) 84 CHAPTER 3. INTEGRATION ON MANIFOLDS where gµν=d/summationdisplay a=1∂xa ∂ξµ∂xa ∂ξν. (3.26) Thevolume form associated with the induced metric is d(Volume) =√gdξ1···dξp, (3.27) whereg= det (gµν). The integral of this p-form over the surface gives the area, orp-dimensional volume, of the surface. If we change the parameterization of the surface from ξµtoζµ, neither thedξ1···dξpnor the√gare separately invariant, but the Jacobean arising from the change of the p-form,dξ1···dξp→dζ1···dζpcancels against the factor coming from the transformation law of the metric tens orgµν→g/prime µν, leading to√gdξ1···dξp=/radicalbig g/primedζ1···dζp. (3.28) The volume of the surface is therefore independent of the co- ordinate system used to evaluate it. Example: The induced metric on the surface of a unit-radius two-spher e embedded in R3, is, expressed in polar angles, “ds2” =g(,) =dθ⊗dθ+ sin2θdφ⊗dφ. Thus g=/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 sin2θ/vextendsingle/vextendsingle/vextendsingle/vextendsingle= sin2θ, and d(Area) = sin θdθdφ. 3.3 Stokes’ Theorem All the integral theorems of classical vector calculus are s pecial cases of Stokes’ Theorem : If∂Ω denotes the (oriented) boundary of the (oriented) region Ω, then /integraldisplay Ωdω=/integraldisplay ∂Ωω. 3.3. STOKES’ THEOREM 85 We will not provide a detailed proof. Apart from notation, it would parallel the proof of Stokes’ or Green’s theorems in ordinar y vector calculus: The exterior derivative dhas been defined so that the theorem holds for an infinitesimal square, cube, or hypercube. We therefore di vide Ω into many such small regions. We then observe that the contributi ons of the interior boundary faces cancel because all interior faces a re shared between two adjacent regions, and so occur twice with opposite orien tations. Only the contribution of the outer boundary remains. Example : If Ω is a region of R2, then from d/bracketleftbigg1 2(xdy−ydx)/bracketrightbigg =dxdy, we have Area(Ω) =/integraldisplay Ωdxdy=1 2/integraldisplay ∂Ω(xdy−ydx). Example : Again, if Ω is a region of R2, then from d[r2dθ/2] =rdrdθ we have Area (Ω) =/integraldisplay Ωrdrdθ =1 2/integraldisplay ∂Ωr2dθ. Example : If Ω is the interior of a sphere of radius R, then /integraldisplay Ωdxdydz =/integraldisplay ∂Ωxdydx =4 3πR3. Here we have referred back to (3.24) to evaluate the surface i ntegral. Example: Archimedes’ tombstone. Archimedes of Syracuse gave instructions that his tombston e should have displayed on it a diagram consisting of a sphere and circumsc ribed cylinder. Cicero, while serving as quæstor in Sicily, had the stone res tored.2This has been said to be the only significant contribution by a Roma n to pure mathematics. The carving on the stone was to commemorate Arc himedes’ results about the areas and volumes of spheres, including th e one illustrated in figure 3.6, that the area of the spherical cap cut off by slici ng through the cylinder is equal to the area cut off on the cylinder. We can understand this result via Stokes’ theorem: If the two -sphereS2 is parameterized by spherical polar co-ordinates θ,φ, and Ω is a region on 2Marcus Tullius Cicero, Tusculan Disputations , Book V, Sections 64 −66 86 CHAPTER 3. INTEGRATION ON MANIFOLDS 1−cos 0 0θθ Figure 3.6: Sphere and circumscribed cylinder. the sphere, then Area (Ω) =/integraldisplay Ωsinθdθdφ =/integraldisplay ∂Ω(1−cosθ)dφ, and applying this to the figure, where the cap is defined by θ<θ 0gives Area (cap) = 2 π(1−cosθ0) which is indeed the area of the blue cylinder. Exercise 3.2 : The sphere Sncan be thought of as the locus of points in Rn+1 obeying/summationtextn+1 i=1(xi)2= 1. Use its invariance under orthogonal transformations to show that the element of surface “volume” of the n-sphere can be written as d(Volume on Sn) =1 n!/epsilon1α1α2...αn+1xα1dxα2...dxαn+1. Use Stokes’ theorem to relate the integral of this form over t hesurface of the sphere to the volume of the solidunit sphere. Confirm that we get the correct proportionality between the volume of the solid unit sphere and the volume or area of its surface. 3.4. APPLICATIONS 87 3.4 Applications We now know how to integrate forms. What sort of forms should w e seek to integrate? For a physicist working with a classical or qua ntum field, a plentiful supply of intesting forms is obtained by using the field to pull back geometric objects. 3.4.1 Pull-backs and Push-forwards If we have a map φfrom a manifold Mto another manifold N, and we choose a pointx∈M, we can push forward a vector from TMxtoTNφ(x), in the obvious way (map head-to-head and tail-to-tail). This map i s denoted by φ∗:TMx→TNφ(x). xx+X XXφ*φ(x)φ(x+X)M N φ Figure 3.7: Pushing forward a vector XfromTMxtoTNφ(x). If the vector Xhas components Xµand the map takes the point with coor- dinatesxµto one with coordinates ξµ(x), the vector φ∗Xhas components (φ∗X)µ=∂ξµ ∂xνXν. (3.29) This looks very like the transformation formula for contrav ariant vector com- ponents under a change of coordinate system. What we are doin g here is conceptually different, however. A change of co-ordinates p roduces a passive transformation — i.e.a new description for an unchanging vector. A push forward is an active transformation — we are changing a vector into differ- ent one. Furthermore, the map from M→Nis not being assumed to be 88 CHAPTER 3. INTEGRATION ON MANIFOLDS one-to-one, so, contrary to the requirement imposed on a co- ordinate trans- formation, it may not be possible to invert the functions ξµ(x) and write the xν’s as functions of the ξµ’s. While we can push forward individual vectors, we cannot alwa ys push forward a vector fieldXfromTMtoTN. If two distinct points x1andx2, chanced to map to the same point ξ∈N, andX(x1)/negationslash=X(x2), we would not know whether to chose φ∗[X(x1)] orφ∗[X(x2)] as [φ∗X](ξ). This problem does not occur for differential forms. A map φ:M→Ninduces a natural, and always well defined, pull-back mapφ∗:/logicalandtextp(T∗N)→/logicalandtextp(T∗M) which works as follows: Given a form ω∈/logicalandtextp(T∗N), we define φ∗ωas a form on M by specifying what we get when we plug the vectors X1,X2,...,Xp∈TM into it. We evaluate the form at x∈Mby pushing the vectors Xi(x) forward fromTMxtoTNφ(x), plugging them into ωatφ(x) and declaring the result to be the evaluation of φ∗ωon theXiatx. Symbolically [φ∗ω](X1,X2,...,Xp) =ω(φ∗X1,φ∗X2,...,φ ∗Xp). (3.30) This may seem rather abstract, but the idea is in practice qui te simple: If the map takes x∈M→ξ(x)∈N, and ω=1 p!ωi1...ip(ξ)dξi1...dξip, (3.31) then φ∗ω=1 p!ωi1i2...ip[ξ(x)]dξi1(x)dξi 2(x)···dξip(x) =1 p!ωi1i2...ip[ξ(x)]∂ξi1 ∂xµ1∂ξi2 ∂xµ2···∂ξip ∂xµ1dxµ1···dxµp.(3.32) Computationally, the process of pulling back a form is so tra nsparent that it easy to confuse it with a simple change of variable. That it is not the same operation will become clear in the next few sections where we consider maps that are many-to-one. Exercise 3.3 : Show that the operation of taking an exterior derivative co m- mutes with a pull back: d[φ∗ω] =φ∗(dω). Exercise 3.4 : If the map φ:M→Nis invertible then we may push forward a vector field XonMto get a vector field φ∗XonN. Show that LX[φ∗ω] =φ∗[Lφ∗Xω]. 3.4. APPLICATIONS 89 Exercise 3.5 : Again assume that φ:M→Nis invertible. By using the co- ordinate expressions for the Lie bracket and the effect of a pu sh-forward, show that ifX,Yare vector fields on TMthen φ∗([X,Y]) = [φ∗X,φ ∗Y], as vector fields on TN. 3.4.2 Spin textures As an application of pull-backs we will consider some of the t opological as- pects of spin textures which are fields of unit vectors n, or “spins”, in two or three dimensions. Consider a smooth map n:R2→S2that assigns x/mapsto→n(x), where nis a three-dimensional unit vector whose tip defines a point on th e 2-sphereS2. A physical example of such an n(x) would be the local direction of the spin polarization in a ferromagnetically-coupled two-dimensi onal electron gas. In terms of n, the area 2-form on the sphere becomes Ω =1 2n·(dn×dn)≡1 2/epsilon1ijknidnjdnk. (3.33) Thenmap pulls this area-form back to F≡n∗Ω =1 2(/epsilon1ijkni∂µnj∂νnk)dxµdxν= (/epsilon1ijkni∂1nj∂2nk)dx1dx2(3.34) which is a differential form in R2. We will call it the topological charge density . It measures the area on the two-sphere swept out by the nvectors as we explore a square in R2of sidedx1bydx2. Suppose now that the vector ntends some fixed direction at large dis- tance. This allows us to think of “infinity” as a single point, and the assign- mentx/mapsto→n(x) as a map from S2toS2. Such maps are characterized topo- logically by their “topological charge,” orwinding number Nwhich counts the number of times the image of the originating xsphere wraps round the target n-sphere. A mathematician would call this number the Brouwer de- greeof the map n. It is intuitively plausible that a continuous map from a sphere to itself will wrap a whole number of times, and so we ex pect N=1 4π/integraldisplay R2/braceleftbig /epsilon1ijkni∂1nj∂2nk/bracerightbig dx1dx2, (3.35) 90 CHAPTER 3. INTEGRATION ON MANIFOLDS to be an integer. We will soon show that this is indeed so, but fi rst we will demonstrate that Nis atopological invariant . In two dimensions the form F=n∗Ω is automatically closed because the exterior derivative of any two-form is zero — there being no three-forms in two dimensions. Even if we consider an n(x1,...,xm) field inm > 2 dimensions, however, we still have dF= 0. This is because dF=1 2/epsilon1ijk∂σni∂µnj∂νnkdxσdxµdxν. (3.36) If we insert infinitesimal vectors into the dxµto get their components δxµ, we have to evaluate the triple-product of three vectors δni=∂µniδxµ, each of which is tangent to the two-sphere. But the tangent space o fS2is two- dimensional, so any three tangent vectors t1,t2,t3, are linearly dependent and their triple-product t1·(t2×t3) is zero. Although it is closed, F=n∗Ω will not generally be the dof a globally defined one-form. Suppose, however, that we vary the map, n→n+δn. The change in the topological charge density is δF=n∗[n·(d(δn)×dn)], (3.37) and this variation canbe written as a total derivative δF=d{n∗[n·(δn×dn)]}≡d{/epsilon1ijkniδnj∂µnkdxµ}. (3.38) In these manipulations we have used δn·(dn×dn) =dn·(δn×dn) = 0, the triple-products being zero for the same reason adduced earl ier. From Stokes’ theorem, we have δN=/integraldisplay S2δF=/integraldisplay ∂S2/epsilon1ijkniδnj∂µnkdxµ. (3.39) Since∂S2=∅, we conclude that δN= 0 under any smooth deformation of the map n(x). This is what we mean when we say that Nis a topological invariant. Equivalently, on R2, with nconstant at infinity, we have δN=/integraldisplay R2δF=/integraldisplay Γ/epsilon1ijkniδnj∂µnkdxµ, (3.40) where Γ is a curve surrounding the origin at large distance. A gainδN= 0, this time because ∂µnk= 0 everywhere on Γ. 3.4. APPLICATIONS 91 In some physical applications, the field nwinds in localized regions called Skyrmions . These knots in the spin field behave very much as elementary particles, retaining their identity as they move through th e material. The winding number counts how many Skyrmions (minus the number o f anti- Skyrmions, which wind with opposite orientation) there are . To construct a smooth multi-Skyrmion map R2→S2with positive winding number N, take a set ofN+ 1 complex numbers λ,a1,...,aNand another set of Nnumbers b1,...,bNsuch that no bcoincides with any a. Then set eiφtanθ 2=λ(z−a1)...(z−aN) (z−b1)...(z−bN)(3.41) wherez=x1+ix2, andθandφare spherical polar co-ordinates specifying the direction n. At the points aithe vector npoints straight up, and at the pointsbiit points straight down. You will show in exercise 3.12 that t his particular n-field configuration minimizes the energy functional E[n] =1 2/integraldisplay (∂1n·∂1n+∂2n·∂2n)dx1dx2 =1 2/integraldisplay/parenleftbig |∇n1|2+|∇n2|2+|∇n3|2/parenrightbig dx1dx2(3.42) for the given winding number N. The next section will explain the geometric origin of the mysterious combination eiφtanθ/2. 3.4.3 The Hopf Map You may recall that in section 1.2.3 we defined complex projective space CPnto be the set of raysin a complex n+ 1 dimensional vector space. A ray is an equivalence classes of vectors [ ζ1,ζ2,...,ζn+1], where the ζiare not all zero, and where we do not distinguish between [ ζ1,ζ2,...,ζn+1] and [λζ1,λζ2,...,λζn+1] for non-zero λ. The space of rays is a 2 n-dimensional real manifold: in a region where ζn+1does not vanish, we can take as co-ordinates the real numbers ξ1,...,ξn,η1,...,ηnwhere ξ1+iη1=ζ1 ζn+1, ξ 2+iη2=ζ2 ζn+1,...,ξn+iηn=ζn ζn+1. (3.43) Similar co-ordinate charts can be constructed in the region s where other ζiare non-zero. Every point in CPnlies in at least one of these co-ordinate charts, 92 CHAPTER 3. INTEGRATION ON MANIFOLDS and the co-ordinate transformation rules for going from cha rt to another are smooth. The simplest complex projective space, CP1, is the real two-sphere S2in disguise. This rather non-obvious fact is revealed by the us e of astereographic mapto make the equivalence class [ ζ1,ζ2]∈CP1correspond to a point non the sphere. When ζ1is non zero, the class [ ζ1,ζ2] is uniquely determined by the ratioζ2/ζ1=|ζ2/ζ1|eiφ, which we plot on the complex plane. We think of this copy of Cas being the x,yplane in R3. We then draw a straight line connecting the plotted point to the south pole of a unit spher e circumscribed about the origin in R3. The point where this line (continued if necessary) intersects the sphere is the tip of the unit vector n. θ θ/2S2 n1ζ /ζ 21=ζN Sz y nζx SN Figure 3.8: Two views of the sterographic map between the two-sphere and the complex plane. The point ζ=ζ2/ζ1∈Ccorresponds to the unit vector n∈S2. Ifζ2, were zero, we would end up at the north pole where the R3co-ordinate ztakes the value z= 1. Ifζ1goes to zero with ζ2fixed, we move smoothly to the south pole z=−1. We therefore extend the definition of our map to the caseζ1= 0 by making the equivalence class [0 ,ζ2] correspond to the south pole. We can find an explicit formula for this map. Figure 3.8 s hows that ζ2/ζ1=eiφtanθ/2, and this relation suggests the use of the “ t”-substitution formulae sinθ=2t 1 +t2,cosθ=1−t2 1 +t2, (3.44) wheret= tanθ/2. Since the x,y,z components of nare given by n1= sinθcosφ, 3.4. APPLICATIONS 93 n2= sinθsinφ, n3= cosθ, we find that n1+in2=2(ζ2/ζ1) 1 +|ζ2/ζ1|2, n3=1−|ζ2/ζ1|2 1 +|ζ2/ζ1|2. (3.45) We can multiply through by |ζ1|2=ζ1ζ1, and so write this correspondence in a more symmetrical manner: n1=ζ1ζ2+ζ2ζ1 |ζ1|2+|ζ2|2 n2=1 i/parenleftbiggζ1ζ2−ζ2ζ1 |ζ1|2+|ζ2|2/parenrightbigg , n3=|ζ1|2−|ζ2|2 |ζ1|2+|ζ2|2. (3.46) This last form can be conveniently expressed in terms of the P auli sigma matrices ˆσ1=/parenleftbigg 0 1 1 0/parenrightbigg ,ˆσ2=/parenleftbigg 0−i i0/parenrightbigg ,ˆσ3=/parenleftbigg 1 0 0−1/parenrightbigg . (3.47) as n1= (z1,z2)/parenleftbigg 0 1 1 0/parenrightbigg/parenleftbigg z1 z2/parenrightbigg , n2= (z1,z2)/parenleftbigg 0−i i0/parenrightbigg/parenleftbigg z1 z2/parenrightbigg , n3= (z1,z2)/parenleftbigg 1 0 0−1/parenrightbigg/parenleftbigg z1 z2/parenrightbigg , (3.48) where /parenleftbigg z1 z2/parenrightbigg =1/radicalbig |ζ1|2+|ζ2|2/parenleftbigg ζ1 ζ2/parenrightbigg (3.49) is a normalized 2-vector, which we think of as a spinor . TheCP1/similarequalS2correspondence now has a quantum mechanical interpre- tation: Any unit three-vector ncan be obtained as the expectation value 94 CHAPTER 3. INTEGRATION ON MANIFOLDS of the ˆσmatrices in a normalized spinor state. Conversly, any norma lized spinorψ= (z1,z2)Tgives rise to a unit vector via ni=ψ†ˆσiψ. (3.50) Now, since 1 =|z1|2+|z2|2, (3.51) the normalized spinor can be thought of as defining a point in S3. This means that the one-to-one correspondence [ z1,z2]↔nalso gives rise to a map fromS3→S2. This is called the Hopf map : Hopf :S3→S2. (3.52) The dimension reduces from three to two, so the Hopf map canno t be one-to- one. Even after we have normalized [ ζ1,ζ2], we are still left with a choice of overall phase. Both ( z1,z2) and (z1eiθ,z2eiθ), although distinct points in S3, correspond to the same point in CP1, and hence in S2. The inverse image of a point in S2is a geodesic circle in S3. Later we will show that any two such geodesic circles are linked, and this makes the Hopf map topologically non-trivial in that it cannot be continuously deformed to a c onstant map, i.e.to a map that takes all of S3to a single point in S2. Exercise 3.6 : We have seen that the stereographic map relates the point wi th spherical polar co-ordinates θ,φto the complex number ζ=eiφtanθ/2. We can therefore set ζ=ξ+iηand takeξ,ηasstereographic co-ordinates on the sphere. Show that in these co-ordinates the sphere metri c is given by g(,)≡dθ⊗dθ+ sin2θdφ⊗dφ =2 (1 +|ζ|2)2(dζ⊗dζ+dζ⊗dζ) =4 (1 +ξ2+|η|2)2(dξ⊗dξ+dη⊗dη), and the area 2-form becomes Ω≡sinθdθ∧dφ =2i (1 +|ζ|2)2dζ∧dζ =4 (1 +ξ2+η2)2dξ∧dη. (3.53) 3.4. APPLICATIONS 95 3.4.4 Homotopy and the Hopf map We can use the Hopf map to factor the map n:x/mapsto→n(x) through the three- sphere by specifying the spinor ψat each point, instead of the vector n, and so mapping indirectly R2ψ→S3Hopf→S2. It might seem that for a given spin-field n(x) we can choose the overall phase ofψ(x)≡(z1(x),z2(x))Tas we like, but if we demand that the zi’s be continuous functions of xthere is a rather non-obvious topological restriction which has important physical consequences. To see how this c omes about we first express the winding number in terms of the zi. We find (after a page or two of algebra) F= (/epsilon1ijkni∂1nj∂2nk)dx1dx2=2 i2/summationdisplay i=1(∂1zi∂2zi−∂2zi∂1zi)dx1dx2,(3.54) and so the topological charge Nis given by N=1 2πi/integraldisplay2/summationdisplay i=1(∂1zi∂2zi−∂2zi∂1zi)dx1dx2. (3.55) Now, when written in terms of the zivariables, the form Fbecomes a total derivative: F=2 i2/summationdisplay i=1(∂1zi∂2zi−∂2zi∂1zi)dx1dx2 =d/braceleftBigg 1 i2/summationdisplay i=1(zi∂µzi−(∂µzi)zi)dxµ/bracerightBigg . (3.56) Further, because nis fixed at large distance, we have ( z1,z2) =eiθ(c1,c2) near infinity, where c1,c2are constants with |c1|2+|c2|2= 1. Thus, near infinity, 1 2i2/summationdisplay i=1(zi∂µzi−(∂µzi)zi)→(|c1|2+|c2|2)dθ=dθ. (3.57) We combine this observation with Stokes’ theorem to obtain N=1 2πi/integraldisplay Γ1 22/summationdisplay i=1(zi∂µzi−(∂µzi)zi)dxµ=1 2π/integraldisplay Γdθ. (3.58) 96 CHAPTER 3. INTEGRATION ON MANIFOLDS Here, as in the previous section, Γ is a curve surrounding the origin at large distance. Now/integraltext dθis the total change in θas we circle the boundary. While the phaseeiθhas to return to its original value after a round trip, the ang le θcan increase by an integer multiple of 2 π. The winding number/contintegraltext dθ/2π can therefore be non-zero, but must be an integer. We have uncovered the rather surpring fact that the topologi cal charge of the map n:S2→S2is equal to the winding number of the phase angle θat infinity. This is the topological constraint refered to ea rlier. As a byproduct, we have confirmed our conjecture that the topolog ical charge N is an integer. The existence of this integer invariant shows that the smooth mapsn:S2→S2fall into distinct homotopy classes labeled by N. Maps with different values of Ncannot be continuously deformed into one another, and, while we have not shown that it is so, two maps with the sam e value of Ncan be deformed into each other. Maps that can be continuously deformed one into the other are said to behomotopic . The set of homotopy classes of the maps of the n-sphere into a manifold Mis denoted by πn(M). In the present case M=S2. We are therefore claiming that π2(S2) =Z, (3.59) where we are identifying the homotopy class with its winding numberN∈Z. 3.4.5 The Hopf index We have so far discussed maps from S2toS2. It is perhaps not too surprising that such maps are classified by a winding number. What is rath er more surprising is that maps n:S3→S2also have an associated topological number. If we continue to assume that ntends to a constant direction at infinity so that we can think of R3∪{∞} as beingS3, this number will label the homotopy classes π3(S2) of fields of unit vectors ninthree dimensions. We will think of the third dimension as being time. In this sit uation an interesting set of nfields to consider are the n(x,t) corresponding moving Skyrmions. The world lines of these Skyrmions will be tubes o utside of which nis constant, and such that on any slice through the tube, nwill cover the target n-sphere once. To motivate the formula we will find for the topological numbe r, we begin with a problem from magnetostatics. Suppose we are given a ca ble originally made up of a bundle of many parallel wires. The cable is then tw istedN 3.4. APPLICATIONS 97 I Figure 3.9: A twisted cable with N= 5. times about its axis and bent into a closed loop, the end of eac h individual wire being attached to its begining to make a continuous circ uit. A current Iflows in the cable in such a manner that each individual wire ca rries only a small part δIiof the total. The sense of the current is such that as we flow with it around the cable each wire wraps Ntimes anticlockwise about all the others. The current produces a magnetic field B. Can we determine the integer twisting number Nknowing only this Bfield? The answer is yes. We use Ampere’s law in integral form, /contintegraldisplay ΓB·dr= (current encircled by Γ) . (3.60) We also observe that the current density ∇×B=Jat a point is directed along the tangent to the wire passing through that point. We t herefore integrate along each individual wire as it encircles the oth ers, and sum over the wires to find /summationdisplay wiresiδIi/contintegraldisplay B·dri=/integraldisplay B·Jd3x=/integraldisplay B·(∇×B)d3x=NI2.(3.61) We now apply this insight to our three-dimensional field of un it vectors n(x). 98 CHAPTER 3. INTEGRATION ON MANIFOLDS The quantity playing the role of the current density Jis the topological cur- rent Jσ=1 2/epsilon1σµν/epsilon1ijkni∂µnj∂νnk. (3.62) We note that∇·J= 0. This is simply another way of saying that the 2-form F=n∗Ω is closed. The flux of Jthrough a surface Sis /integraldisplay SJ·dS=/integraldisplay SF (3.63) and this is the area of the spherical surface covered by the n’s. A Skyrmion, for example, has total topological current I= 4π, the total surface area of the 2-sphere. The Skyrmion world-line will play the role of t he cable, and the inverse images in R3of points on S2correspond to the individual wires. If form language, the field corresponding to Bcan be any one-form A such thatdA=F. Thus NHopf=1 I2/integraldisplay R3B·Jd3x=1 16π2/integraldisplay R3AF (3.64) will be an integer. This integer is the Hopf linking number , orHopf index , and counts the number of times the Skyrmion twists before it b ites its tail to form a closed-loop world-line. There is another way of obtaining this formula, and of unders tanding the number 16π2. We observe that the two-form Fand the one-form Aare the pull-back from S3toR3alongψof the forms F=1 i2/summationdisplay i=1(dzidzi−dzidzi), A=1 i2/summationdisplay i=1(zidzi−zidzi), (3.65) respectively. If we substitute z1,2=ξ1,2+iη1,2, we find that AF= 8(ξ1dη1dξ2dη2−η1dξ1dξ2dη2+ξ2dη2dξ1dη1−η2dξ2dξ1dη1).(3.66) We know from exercise 3.2 that this expression is eight times the volume 3-form on the three-sphere. Now the total volume of the unit t hree-sphere is 2π2, and so, from our factored map x/mapsto→ψ/mapsto→nwe have that NHopf=1 16π2/integraldisplay R3ψ∗(AF) =1 2π2/integraldisplay R3ψ∗d(Volume on S3) (3.67) 3.4. APPLICATIONS 99 is the number of times the normalized spinor ψ(x) coversS3asxcovers R3. For the Hopf map itself, this number is unity, and so the loop i nS3which is the inverse image of a point in S2will twist once around any other such inverse image loop. We have now established that π3(S2) =Z. (3.68) This result, implying that there are many maps from the three -sphere to the two-sphere that are not smoothly deformable to a constan t map, was an great surprise when Hopf discovered it. One of the principal physics consequences of the existence o f the Hopf index is that “quantum lump” quasi-particles like the Skyrm ion can be fermions, even though they are described by commuting (and t herefore bo- son) fields. To understand how this can be, we first explain tha t the collection of homotopy classes πn(M) is not just a set. It has the additional structure of being a group: we can compose two homotopy classes to get a third, the composition is associative, and each homotopy class has an i nverse. To define the group composition law, we think of Snas the interior of an n-dimensional cube with the map f:Sn→Mtaking a fixed value m0∈Mat all points on the boundary of the cube. The boundary can then be consider ed to be a single point on Sn. We then take one of the ndimensions as being “time” and place two cubes and their maps f1,f2into contact, with f1being “ear- lier” andf2being “later.” We thus get a continuous map from a bigger box intoM. The homotopy class of this map, after we relax the condition that the map takes the value m0on the common boundary, defines the composi- tion [f2]◦[f1] of the two homotopy classes corresponding to f1andf2. The composition may be shown to be independent of the choice of re presentative functions in the two classes. The inverse of a homotopy class [f] is obtained by reversing the direction of “time” for each of the maps in th e class. This group structure appears to depend on the fixed point m0. As long as M is arcwise connected, however, the groups obtained from diff erentm0’s are isomorphic , or equivalent. In the case of π2(S2) =Zandπ3(S2) =Z, the composition law is simply the addition of the integers N∈Zthat label the classes. A full account of homotopy theory for working physi cists is to be found in a readable review article by David Mermin.3 3N. D. Mermin, “The topological theory of defects in ordered m edia.” Rev. Mod. Phys. 51(1979) 591. 100 CHAPTER 3. INTEGRATION ON MANIFOLDS When we quantize using Feynman’s “sum over histories” path i ntegral, we may multiply the contributions of histories fthat are not deformable into one another by different phase factors exp {iφ([f])}. The choice of phases must, however, be compatible with the composition of histor ies by concate- nating one after the other – essentially the same operation a s composing homotopy classes. This means that the product exp {iφ([f1]))}exp{iφ([f2])} of the phase factors for two possible histories must be the ph ase factor exp{iφ([f2]◦[f1])}assigned to the composition of their homotopy classes. If our quantum system consists of spins nin two space and one time di- mension we can consistently assign a phase factor exp( iπNHopf) to a history. The rotation of a single Skyrmion through 2 πmakesNHopf= 1 and so the wavefunction changes sign. We will show in the next section, that a his- tory where two particles change places can be continuously d eformed into a history where they do not interchange, but instead one of the m is twisted through 2π. The wavefunction of a pair of Skyrmions therefore changes s ign when they are interchanged. This means that the quantized Sk yrmion is a fermion. 3.4.6 Twist and Writhe Consider two oriented non-intersecting closed curves γ1andγ2. We can use Amp` ere’s law to count the number of times γ1encirclesγ2by imagining that γ2carries a unit current in the direction of its orientation, a nd evaluating Lk(γ1,γ2) =/contintegraldisplay γ1B(r1)·dr1 =1 4π/contintegraldisplay γ1/contintegraldisplay γ2(r1−r2)·(dr1×dr2) |r1−r2|3. (3.69) Here the second line follows from the first by an application o f the Biot-Savart law to compute the Bfield due the current. The second line shows that the Gauss linking number Lk(γ1,γ2) is symmetric under the interchange γ1↔γ2 of the two curves. It changes sign, however, if one of the curv es changes orientation, or if the pair of curves is reflected in a mirror. Introduce parameters t1,t2with 0<t1,t2≤1 to label points on the two curves. The curves are closed, so r1(0) =r1(1), and similarly for r2. Let us also define a unit vector n(t1,t2) =r1(t1)−r2(t2) |r1(t1)−r2(t2)|. (3.70) 3.4. APPLICATIONS 101 Then Lk(γ1,γ2) =1 4π/contintegraldisplay γ1/contintegraldisplay γ2r1(t1)−r2(t2) |r1(t1)−r2(t2)|3·/parenleftbigg∂r1 ∂t1×∂r2 ∂t2/parenrightbigg dt1dt2 =−1 4π/integraldisplay T2n·/parenleftbigg∂n ∂t1×∂n ∂t2/parenrightbigg dt1dt2. (3.71) is seen to be (minus) the winding number of the map n: [0,1]×[0,1]→S2. (3.72) of the 2-torus into the sphere. Our previous results on maps i nto the 2-sphere therefore confirm our Amp` ere-law intuition that Lk( γ1,γ2) is an integer. The linking number is also topological invariant, being unchan ged under any de- formation of the curves that does not cause one to pass throug h the other. An important application of these ideas occurs in biology, w here the curves are the two complementary strands of a closed loop of D NA. We can think of two such parallel curves as forming the edges of a ribbon{γ1,γ2}of width/epsilon1. Let use denote by γthe curve r(t) running along the axis of the ribbon midway between γ1andγ2. The unit tangent to γat the point r(t) is t(t) =˙r(t) |˙r(t)|, (3.73) where the dots denote differentiation with respect to t. We also introduce a unit vector u(t) that is perpendicular to t(t) and lies in the ribbon, pointing fromr1(t) tor2(t). t u γγ12 Figure 3.10: An oriented ribbon {γ1,γ2}showing the vectors tandu. 102 CHAPTER 3. INTEGRATION ON MANIFOLDS We will assign a common value of the parameter tto a point on γand the points nearest to r(t) onγ1andγ2. Consequently r1(t) =r(t)−1 2/epsilon1u(t) r2(t) =r(t) +1 2/epsilon1u(t) (3.74) We can express ˙uas ˙u=ω×u (3.75) for some angular-velocity vector ω(t). The quantity Tw =1 2π/contintegraldisplay γ(ω·t)dt (3.76) is called the Twist of the ribbon. It is not usually an integer, and is a property of the ribbon {γ1,γ2}itself, being independent of the choice of parameterization t. If we set r1(t) andr2(t) equal to the single axis curve r(t) in the integrand of (3.69), the resulting “self-linking” integral, or Writhe , Wrdef=1 4π/contintegraldisplay γ/contintegraldisplay γ(r(t1)−r(t2))·(˙r(t1)×˙r(t2)) |r(t1)−r(t2)|3dt1dt2. (3.77) remains convergent despite the factor of |r(t1)−r(t2)|3in the denominator. However, if we try to achieve this substitution by making the width of the ribbon/epsilon1tend to zero, we find that the vector n(t1,t2) abruptly reverses its direction as t1passest2. In the limit of infinitesimal width this violent motion provides a delta-function contribution −(ω·t)δ(t1−t2)dt1∧dt2 (3.78) to the 2-sphere area swept out by n, and this contribution is invisible to the Writhe integral. The Writhe is a property only of the overall shape of the axis curveγ, and is independent both of the ribbon that contains it, and o f the choice of parameterization. The linking number, on the o ther hand, is independent of /epsilon1, so the/epsilon1→0 limit of the linking-number integral is not the integral of the /epsilon1→0 limit of its integrand. Instead we have Lk(γ1,γ2) =1 2π/contintegraldisplay γ(ω·t)dt+1 4π/contintegraldisplay γ/contintegraldisplay γ(r(t1)−r(t2))·(˙r(t1)×˙r(t2)) |r(t1)−r(t2)|3dt1dt2 (3.79) 3.4. APPLICATIONS 103 This formula Lk = Tw + Wr (3.80) is known as the Calugareanu-White-Fuller relation, and is the basis for the claim, made in the previous section, that the worldline of an extended particle with an exchange (Wr = ±1) can be deformed into a worldline with a 2 π rotation (Tw =±1) without changing the topologically invariant linking number. 1 t2 t11 0t−tΓ Γ tt( ) −tt( ) Figure 3.11: Cutting and reassembling the domain of integration in (3.82). By setting n(t1,t2) =r(t1)−r(t2) |r(t1)−r(t2)|. (3.81) we can express the Writhe as Wr =−1 4π/integraldisplay T2n·/parenleftbigg∂n ∂t1×∂n ∂t2/parenrightbigg dt1dt2, (3.82) but we must take care to recognize that this new n(t1,t2) is discontinuous across the line t=t1=t2. It is equal to t(t) fort1infinitesimally larger thant2, and equal to−t(t) whent1is infinitesimally smaller than t2. By cutting the square domain of integration and reassembling i t into a rhom- boid, as shown in figure 3.11, we obtain a continuous integran d and see that the Writhe is (minus) the 2-sphere area (counted with multip licies and di- vided by 4π) of a region whose boundary is composed of two curves Γ, the tangent indicatrix , ortantrix , on which n=t(t), and its oppositely oriented antipodal counterpart Γ/primeon which n=−t(t). The 2-sphere area Ω(Γ) bounded by Γ is only determined by Γ up t o the addition of integer multiples of 4 π. Taking note that the “wrong” orientation 104 CHAPTER 3. INTEGRATION ON MANIFOLDS of the boundary Γ (see figure 3.11 again) compensates for the m inus sign before the integral in (3.82), we have 4πWr = 2Ω(Γ) + 4 πn. (3.83) Thus, Wr =1 2πΩ(Γ),mod 1. (3.84) We can do better than (3.84) once we realize that by allowing c rossings we can continuously deform any closed curve into a perfect circ le. Each self- crossing causes Lk and Wr (but not Tw which, being a local func tional, does not care about crossings) to jump by ±2. For a perfect circle Wr = 0 whilst Ω = 2π. We therefore have an improved estimate of the additive inte ger that is left undetermined by Γ, and from it we obtain Wr = 1 +1 2πΩ(Γ),mod 2. (3.85) This result is due to Brock Fuller.4 We can use our ribbon language to describe conformational tr ansitions in long molecules. The elastic energy of a closed rod (or DNA mol ecule) can be approximated by E=/integraldisplay γ/braceleftbigg1 2α(ω·t)2+1 2βκ2/bracerightbigg ds (3.86) Here we are parameterizing the curve by its arc-length s. The constant αis the torsional stiffness coefficient, βis the flexural stiffness, and κ(s) =/vextendsingle/vextendsingle/vextendsingle/vextendsingled2r(s) ds2/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingledt(s) ds/vextendsingle/vextendsingle/vextendsingle/vextendsingle, (3.87) is the local curvature. Suppose that our molecule has linkin g numbern,i.e it was twisted ntimes before the ends were joined together to make a loop. 4F. Brock Fuller, Proc. Natl. Acad. Sci. USA, 75(1978) 3557 - 61. 3.5. EXERCISES AND PROBLEMS 105 Figure 3.12: A molecule initially with Lk = 3 ,Tw = 3 ,Wr = 0 writhes to a new configuration with Lk = 3 ,Tw = 0 ,Wr = 3 . Whenβ/greatermuchαthe molecule will minimize its bending energy by forming a planar circle with Wr ≈0 and Tw≈n. If we increase α, or decrease β, there will come a point at which the molecule will seek to save torsi onal energy at the expense of bending, and will suddenly writhe into a new co nfiguration with Wr≈nand Tw≈0. Such twist-to-writhe transformations will be familiar to anyone who has struggled to coil a garden hose or e lectric cable. 3.5 Exercises and Problems Exercise 3.7 :Old exam problem . A two-form is expressed in Cartesian coor- dinates as, ω=1 r3(zdxdy +xdydz +ydzdx ) wherer=/radicalbig x2+y2+z2. a) Evaluate dωforr/negationslash= 0. b) Evaluate the integral Φ =/integraldisplay Pω over the infinite plane P={−∞<x<∞,−∞<y<∞,z= 1}. c) A sphere is embedded into R3by the map ϕ, which takes the point (θ,φ)∈S2to the point ( x,y,z)∈R3, where x=Rcosφsinθ y=Rsinφsinθ z=Rcosθ. Pull backωand find the 2-form ϕ∗ωon the sphere. ( Hint: The form ϕ∗ωis both familiar and simple. If you end up with an intractable mess of trigonometric functions, you have made an algebraic erro r.) 106 CHAPTER 3. INTEGRATION ON MANIFOLDS d) By exploiting the result of part c), or otherwise, evaluat e the integral Φ =/integraldisplay S2(R)ω whereS2(R) is the surface of a two-sphere of radius Rcentered at the origin. The following four exercises all explore the same geometric facts relating to Stokes’ theorem and the area 2-form of a sphere, but in differe nt physical settings. Exercise 3.8 : A flywheel of moment of inertia Ican rotate without friction about an axle whose direction is specified by a unit vector n. The flywheel and axle are initially stationary. The direction nof the axle is made to describe a simple closed curve γ=∂Ω on the unit sphere, and is then left stationary. γΩn Figure 3.13: Flywheel Show that once the axle has returned to rest in its initial dir ection, the flywheel has also returned to rest, but has rotated through an angle θ= Area(Ω) when compared with its initial orientation. The area of Ω is t o be counted as positive if the path γsurrounds it in a clockwise sense, and negative otherwise. Observe that the path γbounds two regions with opposite orientations. Taking into account that we cannot define the rotation angle at inter mediate steps, show that the area of either region can be used to compute θ, the results being physically indistinguishable. (Hint: Show that the c omponentLZ= I(˙ψ+˙φcosθ) of the flywheel’s angular momentum along the axle is a consta nt of the motion.) 3.5. EXERCISES AND PROBLEMS 107 Exercise 3.9 : A ball of unit radius rolls without slipping on a table. The b all moves in such a way that the point in contact with table descri bes a closed pathγ=∂Ω on the ball. (The corresponding path on the table will not necessarily be closed.) Show that the final orientation of th e ball will be such that it has rotated, when compared with its initial orientat ion, through an angleφ= Area(Ω) about a vertical axis through its center, As in the p revious problem, the area is counted positive if γencircles Ω in an anti-clockwise sense. (Hint: recall the no-slip rolling condition ˙φ+˙ψcosθ= 0 from (2.26).) Exercise 3.10 : Let a curve in R3be parameterized by its arc length sasr(s). Then the unit tangent to the curve is given by t(s) =˙r≡dr ds. Theprincipal normal n(s) and the binormal b(s) are defined by the require- ment that ˙t=κnwith the curvatureκ(s) positive, and that t,nandb=t×n form a right-handed orthonormal frame. t b nnbt Figure 3.14: Serret-Frenet frames. a) Show that there exists a scalar τ(s), the torsion of the curve, such that t,nandbobey the Serret-Frenet relations  ˙t ˙n ˙b = 0κ0 −κ0τ 0−τ0  t n b . b) Any pair of mutually orthogonal unit vectors e1(s),e2(s) perpendicular totand such that e1×e2=tcan serve as an orthonormal frame for vectors in the normal plane. A basis pair e1,e2with the property ˙e1·e2−˙e2·e1= 0 108 CHAPTER 3. INTEGRATION ON MANIFOLDS is said to be parallel , orFermi-Walker , transported along the curve. In other words, a parallel-transported 3-frame t,e1,e2slides along the curver(s) in such a way that the component of its angular velocity in thetdirection is always zero. Show that the Serret-Frenet frame e1=n, e2=bisnotparallel transported, but instead rotates at angular veloc ity ˙θ=τwith respect to a parallel-transported frame. c) Consider a finite segment of curve such that the initial and final Serret- Frenet frames are parallel, and so t(s) defines a closed path γ=∂Ω on the unit sphere. Fill in the line-by-line justications fo r the following sequence of manipulations: /integraldisplay γτds =1 2/integraldisplay γ(b·˙n−n·˙b)ds =1 2/integraldisplay γ(b·dn−n·db) =1 2/integraldisplay Ω(db·dn−dn·db) (∗) =1 2/integraldisplay Ω{(db·t)(t·dn)−(dn·t)(t·db)} =1 2/integraldisplay Ω{(b·dt)(dt·n)−(n·dt)(dt·b)} =−1 2/integraldisplay Ωt·(dt×dt) =−Area(Ω). (The line marked ‘ ∗’ is the one that requires most thought. How can we define “ b” and “ n” in the interior of Ω?) d) Conclude that a Fermi-Walker transported frame will have rotated through an angleθ= Area(Ω), compared to its initial orientation, by the time i t reaches the end of the curve. The plane of transversely polarized light propagating in a m onomode optical fibre is Fermi-Walker transported, and this rotation can be s tudied experimen- tally.5 Exercise 3.11 :Foucault’s pendulum (in disguise). A particle of mass mis constrained by a pair of frictionless plates to move in a plan e Π that passes through the origin O. The particle is attracted to O by a force −κr, and it therefore executes simple harmonic motion within Π. The ori entation of the 5A. Tomita, R. Y. Chao, Phys. Rev. Lett. 57(1986) 937-940. 3.5. EXERCISES AND PROBLEMS 109 plane, specified by a normal vector n, can be altered in such a way that Π continues to pass through the centre of attraction O. a) Show that the constrained motion is described by the equat ion m¨r+κr=λ(t)n, and determine λ(t) in terms of m,nand¨r. b) Seek a solution in the form r(t) =A(t)cos(ωt+φ), and, by assuming that nchanges direction slowly compared to the fre- quencyω=/radicalbig κ/m, show that ˙A=−n(˙n·A). Deduce that|A|remains constant, and so ˙A=ω×Afor some angular velocity vector ω. Show thatωis perpendicular to n. c) Show that the results of part b) imply that the direction of oscillation A is “parallel transported” in the sense of the previous probl em. Conclude that if nslowly describes a closed loop γ=∂Ω on the unit sphere, then the direction of oscillation Aends up rotated through an angle θ= Area(Ω). The next exercise introduces an clever trick for solving som e of the non-linear partial differential equations of field theory. The class of e quations to which it and its generalizations are applicable is rather restric ted, but when they work they provide a complete multi-soliton solution. Problem 3.12 : In this problem you will find the spin field n(x) that minimizes the energy functional E[n] =1 2/integraldisplay R2/parenleftbig |∇n1|2+|∇n2|2+|∇n3|2/parenrightbig dx1dx2 for a given positive winding number N. a) Use the results of exercise 3.6 to write the winding number N, defined in (3.35), and the energy functional E[n] as 4πN=/integraldisplay4 (1 +ξ2+η2)2(∂1ξ∂2η−∂1η∂2ξ)dx1dx2, E[n] =1 2/integraldisplay4 (1 +ξ2+η2)2/parenleftbig (∂1ξ)2+ (∂2ξ)2+ (∂1η)2+ (∂2η)2/parenrightbig dx1dx2, whereξandηare stereographic co-ordinates on S2specifying the direc- tion of the unit vector n. 110 CHAPTER 3. INTEGRATION ON MANIFOLDS b) Deduce the inequality E−4πN≡1 2/integraldisplay4 (1 +ξ2+η2)2|(∂1+i∂2)(ξ+iη)|2dx1dx2>0. c) Deduce that for winding number N >0 the minimum energy solutions have energy E= 4πNand are obtained by solving the first-order linear equation/parenleftbigg∂ ∂x1+i∂ ∂x2/parenrightbigg (ξ+iη) = 0. d) Solve the equation in part c) and show that the minimal ener gy solutions with winding number N >0 are given by ξ+iη=λ(z−a1)...(z−aN) (z−b1)...(z−bN) wherez=x1+ix2, andλ,a1,...,aN, andb1,...,bN, are arbitrary complex numbers—except that no amay coincide with any b. This is the solution we displayed at the end of section 3.4.2. e) Repeat the analysis for N < 0. Show that the solutions are given in terms of rational functions of ¯ z=x1−ix2. The idea of combining the energy functional and the topologi cal charge into a single, manifestly positive, functional is due to Evgueny B ogomol’nyi. The the resulting first order linear equation is therefore called a Bogomolnyi equation . If we had tried to find a solution directly in terms of n, we would have ended up with a horribly non-linear second-order partial differen tial equation.. Exercise 3.13 :Lobachevski space . The hyperbolic plane of Lobachevski ge- ometry can be realized by embedding the Z≥Rbranch of the two-sheeted hyperboloid Z2−X2−Y2=R2into a Minkowski space with metric ds2= −dZ2+dX2+dY2. We can parametrize the emebedded surface by making an “imagi nary radius” version of the stereographic map, in which the point P on the h yperboloid is labelled by the co-ordinates of the point Q on the X-Yplane (see figure 3.15). i) Show that the embedding induces the metric g(,) =4R4 (R2−X2−Y2)2(dX⊗dX+dY⊗dY), X2+Y2<R2 of the Poincar´ e disc model (see problem ??.??) on the hyperboloid. 3.5. EXERCISES AND PROBLEMS 111 P QXZ R −R Figure 3.15: A slice through the embedding of two-dimensional Lobachevs ki space into three-dimensional Minkowski space, showing the sterographic pa- rameterization of the embedded space by the Poincar´ e disc X2+Y2<R2. ii) Use the induced metric to show that the area of a disc of hyp erbolic radiusρis given by Area = 4πR2sinh2/parenleftBigρ 2R/parenrightBig = 2πR2(cosh(ρ/R)−1), and so is only given by πρ2whenρis small compared to the scale Rof the hyperbolic space. It suffices to consider circles with the ir centres at the origin. You will first need to show that the hyperbolic dis tanceρ from the center of the disc to a point at Euclidean distance ris ρ=Rln/parenleftbiggR+r R−r/parenrightbigg . Exercise 3.14 : Faraday’s “flux rule” for computing the electromotive forc eE in a circuit containing a thin moving wire is usually derived by the following manipulations: E ≡/contintegraldisplay ∂Ω(E+v×B)·dr =/integraldisplay ΩcurlE·dS−/contintegraldisplay ∂ΩB·(v×dr) =−/integraldisplay Ω∂B ∂t·dS−/contintegraldisplay ∂ΩB·(v×dr) =−d dt/integraldisplay ΩB·dS. 112 CHAPTER 3. INTEGRATION ON MANIFOLDS a) Show that if we parameterize the surface Ω as xµ(u,v,τ ), withu,vla- belling points on Ω and τparametrizing the evolution of Ω, then the corresponding manipulations in the covariant differential -form version of Maxwell’s equations lead to d dτ/integraldisplay ΩF=/integraldisplay ΩLVF=/integraldisplay ∂ΩiVF=−/integraldisplay ∂Ωf whereVµ=∂xµ/∂τandf=−iVF. b) Show that if we take τto be the proper time along the world-line of each element of Ω, then Vis the 4-velocity Vµ=1√ 1−v2(1,v), andf=−iVFbecomes the one-form corresponding to the Lorentz-force 4-vector. It is not clear that the terms in this covariant form of Farday ’s law can be given any physical interpretation outside the low-velocit y limit. When parts of∂Ω have different velocities, the relation of the integrals to measurements made at fixed co-ordinate time requires thought.6 The next pair of exercises explores some physics appearance s of the contin- uum Hopf linking number (3.64). Exercise 3.15 : The equations governing the motion of an incompressible in - viscid fluid are∇·v= 0 and Euler’s equation Dv Dt≡∂v ∂t+ (v·∇)v=−∇P. Recall that the operator ∂/∂t+v·∇, here written as D/Dt , is called the convective derivative . a) Take the curl of Euler’s equation to show that if ω=∇×vis thevorticity thenDω Dt≡∂ω ∂t+ (v·∇)ω= (ω·∇)v. b) Combine Euler’s equation with part a) to show that D Dt(v·ω) =∇·/braceleftbigg ω/parenleftbigg1 2v2−P/parenrightbigg/bracerightbigg . 6See E. Marx, Journal of the Franklin Institute, 300(1975) 353-364. 3.5. EXERCISES AND PROBLEMS 113 c) Show that if Ω is a volume moving with the fluid, then d dt/integraldisplay Ωf(r,t)dV=/integraldisplay ΩDf DtdV. e) Conclude that when ωis zero at infinity the helicity I=/integraldisplay v·(∇×v)dV=/integraldisplay v·ωdV is a constant of the motion. The helicity measures the Hopf linking number of the vortex l ines. The dis- covery7of its conservation founded the field of topological fluid dynamics . Exercise 3.16 : LetB=∇×AandE=−∂A/∂t−∇φbe the electric and magnetic field in an incompressible and perfectly conductin g fluid. In such a fluid the co-moving electromotive force E+v×Bmust vanish everywhere. a) Use Maxwell’s equations to show that ∂A ∂t=v×(∇×A)−∇φ, ∂B ∂t=∇×(v×B). b) From part a) show that the convective derivative of A·Bis given by D Dt(A·B) =∇·{B(A·v−φ)}. c) By using the same reasoning as the previous problem, and as suming that Bis zero at infinity, conclude that Woltjer’s invariant I=/integraldisplay (A·B)dV=/integraldisplay /epsilon1ijkAi∂jAkd3x=/integraldisplay AF is a constant of the motion. This result shows that the Hopf linking number of the magneti c field lines is independent of time. It is an essential ingredient in the geo dynamo theory of the Earth’s magnetic field. 7H. K. Moffatt, J. Fluid Mech. 35(1969) 117. 114 CHAPTER 3. INTEGRATION ON MANIFOLDS Chapter 4 An Introduction to Topology Topology is the study of the consequences of continuity. We a ll know that a continuous real function defined on a connected interval an d positive at one point and negative at another must take the value zero at s ome point between. This fact seems obvious—although a course of real a nalysis will convince you of the need for a proof. A less obvious fact, but o ne that follows from the previous one, is that a continuous function defined on the unit circle must posses two diametrically opposite points a t which it takes the same value. To see that this is so, consider f(θ+π)−f(θ). This difference (if not initially zero, in which case there is nothing furthe r to prove) changes sign asθis advanced through π, because the two terms exchange roles. It was therefore zero somewhere. This observation has practical a pplication in daily life: Our local coffee shop contains four-legged tables that wobble because the floor is not level. They are round tables, however, and bec ause they possess no misguided levelling screws all four legs have the same length. We are therefore guaranteed that by rotating the table about it s center through an angle of less than π/2 we will find a stable location. A ninety-degree rotation interchanges the pair of legs that are both on the gr ound with the pair that are rocking, and at the change-over point all four l egs must be simultaneously on the ground. Similar effects with a practical significance for physics app ear when we try to extend our vector and tensor calculus from a local regi on to an entire manifold. A smooth field of vectors tangent to the sphere S2will always possess a zero — i.e.a point at which the the vector field vanishes. On the torusT2, however, we can construct a nowhere-zero vector field. This shows that the global topology of the manifold influences the way in which 115 116 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY the tangent spaces are glued together to form the tangent bun dle. To study this influence in a systematic manner we need first to understa nd how to characterize the global structure of a manifold, and then to see how this structure affects the mathematical and physical objects tha t live on it. 4.1 Homeomorphism and Diffeomorphism In the previous chapter we met with a number of topological invariants , quantities that are unaffected by continuous deformations. Some invariants help to distinguish topologically distinct manifolds. An i mportant example is the set of Betti numbers of the manifold. If two manifolds have different Betti numbers they are certainly distinct. If, however, they have the same Betti numbers, we cannot be sure that they are topologically ident ical. It is a holy grail of topology to find a complete set of invariants such tha t having them all coincide would be enough to say that two manifolds were to pologically the same. In the previous paragraph we were deliberately vague in our u se of the terms “distinct” and the “same”. Two topological spaces (sp aces equipped with a definition of what is to be considered an open set) are re garded as be- ing the “same”, or homeomorphic , if there is a one-to-one, onto, continuous map between them whose inverse is also continuous. Manifold s come with the additional structure of differentiability: we may therefor e talk of “smooth” maps, meaning that their expression in coordinates is infini tely (C∞) differ- entiable. We regard two manifolds as being the “same”, or diffeomorphic , if there is a one-to-one onto C∞map between them whose inverse is also C∞. The distinction between homeomorphism and diffeomorphism s ounds like a mere technical nicety, but it has consequences for physics. Edward Witten discovered1that there are 992 distinct 11-spheres. These are manifolds that are all homeomorphic to the 11-sphere, but diffeomorphicall y inequivalent. This fact is crucial for the cancellation of global graviati onal anomalies in the E 8×E8or SO(32) symmetric superstring theories. Since we are interested in the consequences of topology for c alculus, we will restrict ourselves to the interpretation “same” = diffe omorphic. 1E. Witten, Comm. Math. Phys. 117(1986), 197. 4.2. COHOMOLOGY 117 4.2 Cohomology Betti numbers arise in answer to what seems like a simple calc ulus problem: when can a vector field whose divergence vanishes be written a s the curl of something? We will see that the answer depends on the global s tructure of the space the field inhabits. 4.2.1 Retractable Spaces: Converse of Poincar´ e Lemma Poincar´ e’s lemma asserts that d2= 0. In traditional vector calculus language this reduces to the statements curl (grad φ) = 0 and div (curl w) = 0. We often assume that the converse is true: If curl v= 0, we expect that we can find aφsuch that v= gradφ, and, if div v= 0, that we can find a wsuch thatv= curl w. You know a formula for the first case: φ(x) =/integraldisplayx x0v·dx, (4.1) but probably do not know the corresponding formula for w. Using differ- ential forms, and provided the space in which these forms liv e has suitable topological properties, it is straightforward to find a solution for the g eneral problem: If ωis closed, meaning that dω= 0, findχsuch thatω=dχ. The “suitable topological properties” referred to in the pr evious para- graph is that the space be retractable . Suppose that the closed form ωis defined in a domain Ω. We say that Ω is retractable to the point O if there exists a smooth map ϕt: Ω→Ω which depends continuously on a parameter t∈[0,1] and for which ϕ1(x) =xandϕ0(x) = O. Applying this retraction map to the form, we will then have ϕ∗ 1ω=ωandϕ∗ 0ω= 0. Let us set ϕt(xµ) =xµ(t). Defineη(x,t) to be the velocity-vector field that corresponds to the co-ordinate flow:dxµ dt=ηµ(x,t). (4.2) An easy exercise, using the interpretation of the Lie deriva tive in (2.40), shows thatd dt(ϕ∗ tω) =Lη(ϕ∗ tω). (4.3) We now use the infinitesimal homotopy relation and our assump tion that dω= 0, and hence (from exercise 3.3) that d(ϕ∗ tω) = 0, to write Lη(ϕ∗ tω) = (iηd+diη)(ϕ∗ tω) =d[iη(ϕ∗ tω)]. (4.4) 118 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY Using this we can integrate up with respect to tto find ω=ϕ∗ 1ω−ϕ∗ 0ω=d/parenleftbigg/integraldisplay1 0iη(ϕ∗ tω)dt/parenrightbigg . (4.5) Thus χ=/integraldisplay1 0iη(ϕ∗ tω)dt, (4.6) solves our problem. This magic formula for χmakes use of the nearly all the “calculus on manifolds” concepts that we have introduced so far. The nota tion is so pow- erful that it has suppressed nearly everything that a tradit ionally-educated physicist would find familiar. We will therefore unpack the s ymbols by means of a concrete example. Let us take Ω to be the whole of R3. This can be retracted to the origin via the map ϕt(xµ) =xµ(t) =txµ. The velocity field whose flow gives xµ(t) =txµ(0) isηµ(x,t) =xµ/t. To verify this, compute dxµ(t) dt=xµ(0) =1 txµ(t), soxµ(t) is indeed the solution to dxµ dt=ηµ(x(t),t). Now let us apply this retraction to ω=Adydz +Bdzdx +Cdxdy with dω=/parenleftbigg∂A ∂x+∂B ∂y+∂C ∂z/parenrightbigg dxdydz = 0. (4.7) The pull-back ϕ∗ tgives ϕ∗ tω=A(tx,ty,tz )d(ty)d(tz) + (two similar terms) . (4.8) The interior product with η=1 t/parenleftbigg x∂ ∂x+y∂ ∂y+z∂ ∂z/parenrightbigg (4.9) 4.2. COHOMOLOGY 119 then gives iηϕ∗ tω=tA(tx,ty,tz )(ydz−zdy) + (two similar terms) . (4.10) Finally we form the ordinary integral over tto get χ=/integraldisplay1 0iη(ϕ∗ tω)dt =/bracketleftbigg/integraldisplay1 0A(tx,ty,tz )tdt/bracketrightbigg (ydz−zdy) +/bracketleftbigg/integraldisplay1 0B(tx,ty,tz )tdt/bracketrightbigg (zdx−xdz) +/bracketleftbigg/integraldisplay1 0C(tx,ty,tz )tdt/bracketrightbigg (xdy−ydx). (4.11) In this expression the integrals in the square brackets are j ust numerical coefficients, i.e., the “dt” is not part of the one-form. It is instructive, because not entirely trivial, to let “ d” act onχand verify that the con- struction works. If we focus first on the term involving A, we find that d[/integraltext1 0A(tx,ty,tz )tdt](ydz−zdy) can be grouped as /bracketleftbigg/integraldisplay1 0/braceleftbigg 2tA+t2/parenleftbigg x∂A ∂x+y∂A ∂y+z∂A ∂z/parenrightbigg/bracerightbigg dt/bracketrightbigg dydz −/integraldisplay1 0t2∂A ∂xdt(xdydz +ydzdx +zdxdy ). (4.12) The first of these terms is equal to /bracketleftbigg/integraldisplay1 0d dt/braceleftbig t2A(tx,ty,tz )/bracerightbig dt/bracketrightbigg dydz=A(x,y,x )dydz, (4.13) which is part of ω. The second term will combine with the terms involving B,C, to become −/integraldisplay1 0t2/parenleftbigg∂A ∂x+∂B ∂y+∂C ∂z/parenrightbigg dt(xdydz +ydzdx +zdxdy ), (4.14) which is zero by our hypothesis. Putting togther the A,B,C, terms does therefore reconstitute ω. 120 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY 4.2.2 Obstructions to Exactness The condition that Ω be retractable plays an essential role i n the converse to Poincar´ e’s lemma. In its absence dω= 0 does not guarantee that there is an χsuch thatω=dχ. Consider, for example, a vector field vwith curl v≡0 in an annulus Ω = {R0<|r|<R 1}. In the annulus (a non-retractable space) the condition that curl v≡0 does not prohibit/contintegraltext Γv·drbeing non zero for some closed path Γ encircling the central hole. When this lin e integral is non-zero then there can be no single-valued χsuch that v=∇χ. If there were such a χ, then /contintegraldisplay Γv·dr=χ(0)−χ(0) = 0. (4.15) A non-zero value for/contintegraltext Γv·drtherefore consititutes an obstruction to the existence of an φsuch that v=∇χ. Example : The sphere S2is not retractable. The area 2-form sin θdθdφ is closed, but, although we can write sinθdθdφ =d[(1−cosθ)dφ], (4.16) the 1-form (1−cosθ)dφis singular at the south pole, θ=π. We could try sinθdθdφ =d[(−1−cosθ)dφ], (4.17) but this is singular at the north pole, θ= 0. There is no escape: we know that /integraldisplay S2sinθdθdφ = 4π, (4.18) but if sinθdθdφ =dχthen Stokes says that /integraldisplay S2sinθdθdφ?=/integraldisplay ∂S2χ= 0 (4.19) because∂S2= 0. Again, a non-zero value for/integraltext ωover some boundary-less region has provided an obstruction to finding an χsuch thatω=dχ. 4.2.3 De Rham Cohomology We have seen that sometimes the condition dω= 0 allows us to find an χsuch thatω=dχ, and sometimes it does not. If the region in which we seek χis 4.2. COHOMOLOGY 121 retractable, we can always construct it. If the region is not retractable there may be an obstruction to the existence of χ. In order to describe the various possibilities we introduce the language of cohomology , or more precisely de Rham cohomology , named for the Swiss mathematician Georges de Rham who did the most to create it. The significance of cohomology for physics is that many impor tant quan- tities can be expressed as integrals of differential forms th at lie in some co- homology space. For simplicity suppose that we are working in a compact manif oldM without boundary. Let Ωp(M) =/logicalandtextp(T∗M) be the space of all smooth p-form fields. It is a vector space over R: we can add p-form fields and multiply them by real constants, but, as is the vector space C∞(M) of smooth functions on M, it is infinite dimensional. The subspace Zp(M) ofclosed forms—those withdω= 0—is also an infinite dimensional vector space, and the same is true of the space Bp(M) ofexact forms — those that can be written as ω=dχfor some globally defined ( p−1)-formχ. Now consider the space Hp=Zp/Bp, which is the space of closed forms modulo exact forms. In this space we do not distinguish between two forms, ω1andω2when there an χ, such thatω1=ω2+dχ. We say that ω1andω2arecohomologous , and write ω1∼ω2∈Hp(M). We will use the symbol [ ω] to denote the equivalence class of forms cohomologous to ω. Now a miracle happens! For a compact manifoldMthe spaceHp(M) isfinite dimensional! It is called the p-th (de Rham) cohomology space of the manifold, and depends only on t he global topology of M. In particular, it does not depend on any metric we may have chosen forM. Sometimes we write Hp DR(M,R) to make clear that we are dealing with de Rham cohomolgy, and that we are working with vector spaces over the real numbers. This is because there is also a space Hp DR(M,Z), where we only allow multiplication by integers. The cohomology space Hp DR(M,R) codifies all potential obstructions to solving the problem of finding a ( p−1)-formχsuch thatdχ=ω: we can find such a χif and only if ωis cohomologous to zero in Hp DR(M,R). If Hp DR(M,R) ={0}, which is the case if Mis retractable, then all closed p- forms are cohomologous to zero. If Hp DR(M,R)/negationslash={0}, then some closed p-formsωwill not be cohomologous to zero. We can test whether ω∼0∈ Hp DR(M,R) by forming suitable integrals. 122 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY 4.3 Homology The language of cohomology seems rather abstract. To unders tand its origin it may be more intuitive to think about the spaces that are the cohomology spaces’ vector-space duals. These homology spaces are simple to understand pictorially. The basic idea is that, given a region Ω, we can find its boundar y∂Ω. Inspection of a few simple cases will soon lead to the conclus ion that the “boundary of a boundary” consists of nothing. In symbols, ∂2= 0. The statement “ ∂2= 0” is clearly analgous to “ d2= 0,” and, pursuing the anal- ogy, we can construct a vector space of “regions” and define tw o “regions” as being homologous if they differ by the boundary of another “region.” 4.3.1 Chains, Cycles and Boundaries We begin by making precise the vague notions of region and bou ndary. Simplicial Complexes The set of all curves and surfaces in a manifold Mis infinite dimensional, but the homology spaces are finite dimensional. Life would be muc h easier if we could use finite dimensional spaces throughout. Mathematic ians therefore do what any computationally-minded physicist would do: the y approximate the smooth manifold by a discrete polygonal grid . Were they i nterested in distances, they would necessarily use many small polygons s o as to obtain a good approximation to the detailed shape of the manifold. T he global topology, though, can often be captured by a rather coarse di scretization. The result of this process is to reduce a complicated problem in differential geometry to one of simple algebra. The resulting theory is th erefore known asalgebraic topology. It turns out to be convenient to approximate the manifold by g eneralized triangles. We therefore dissect Minto line segments (if one dimensional), triangles, (if two dimensional), tetrahedra (if three dime nsional) or higher dimensional p-simplices (singular: simplex ). The rules for the dissection are: a) Every point must belong to at least one simplex. b) A point can belong to only a finite number of simplices. c) Two different simplices either have no points in common, or i) one is a face (or edge, or vertex) of the other, 4.3. HOMOLOGY 123 a) b) Figure 4.1: Triangles, or 2-simplices, that are a) allowed, b) not allow ed in a dissection. In b) only parts of edges are in common. β βP P P Pα αγ a) b)21 γβ P α1 2 Figure 4.2: A triangulation of the 2-torus. a) The torus as a rectangle with periodic boundary conditions: The two edges labled αwill be glued togther point-by-point along the arrows when we reassemble the torus, and so are to be regarded as a single edge. The two sides labeled βwill be glued similarly. b) The assembled torus: All four P’s are now in the same place, and correspond to a single point. ii) the set of points in common is the whole of a shared face (or edge, or vertex). The collection of simplices composing the dissected space i s called a simplicial complex . We will denote it by S. We may not need many triangles to capture the global topology . For example, figure 4.2 shows how a two-dimensional torus can be d ecomposed into two 2-simplices (triangles) bounded by three 1-simpli ces (edges) α,β,γ , and with only a single 0-simplex (vertex) P. Computations are easier to describe, however, if each simplex in the decomposition is u niquely specified by its vertices. For this we usually need a slightly finer diss ection. Figure 4.3 shows a decomposition of the torus into 18 triangles each of which is 124 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY P1 P2 P PP3P1 4P4P P P P P1 P P2 3P15 6 78 9P7 Figure 4.3: A second triangulation of the 2-torus. 1 234 PP P P Figure 4.4: A tetrahedral triangulation of the 2-sphere. The circulati ng arrows on the faces indicate the choice of orientation P1P2P4andP2P3P4. uniquely labeled by three points drawn from a set of nine vert ices. In this figure vertices with identical labels are to be regarded as th e same vertex, as are the corresponding sides of triangles. Thus, each of th e edgesP1P2, P2P3,P3P1, at the top of the figure are to be glued point-by-point to the corresponding edges on bottom of the figure. Similarly along the sides. The resulting simplicial complex then has 27 edges. We may triangulate the sphere S2as a tetrahedron with vertices P1,P2, P3,P4. This dissection has six edges: P1P2,P1P3,P1P4,P2P3,P2P4,P3P4, and four faces: P2P3P4,P1P3P4,P1P2P4andP1P2P3. 4.3. HOMOLOGY 125 p-Chains We assign to simplices an orientation defined by the order in w hich we write their defining vertices. The interchange of of any pair of ver tices reverses the orientation, and we consider there to be a relative minus sig n between oppo- sitely oriented but otherwise identical simplices: P2P1P3P4=−P1P2P3P4. We now construct abstract vector spaces Cp(S,R) ofp-chains which have the oriented p-simplices as their basis vectors. The most general element s of C2(S,R), withSbeing the tetrahedral triangulation of the sphere S2, would be a1P2P3P4+a2P1P3P4+a3P1P2P4+a4P1P2P3, (4.20) wherea1,...,a 4, are real numbers. We regard the distinct faces as being linearly independent basis elements for C2(S,R). The space is therefore four dimensional. If we had triangulated the sphere so that it had 16 triangular faces, the space C2would be 16 dimensional. Similarly, the general element of C1(S,R) would be b1P1P2+b2P1P3+b3P1P4+b4P2P3+b5P2P4+b6P3P4, (4.21) and soC1(S,R) is a six-dimensional space spanned by the edges of the tetra- hedron. For C0(S,R) we have c1P1+c2P2+c3P3+c4P4, (4.22) and soC0(S,R) is four dimensional, and spanned by the vertices . Our manifold comprises only the surface of the two-sphere, so there is no such thing as C3(S,R). The reason for making the field Rexplicit in these definitions is that we sometimes gain more information about the topology if we all ow only integer coefficients. The space of such p-chains is then denoted by Cp(S,Z). Be- cause a vector space requires that coefficients be drawn from a field, these objects are no longer vector spaces. They can be thought of as either mod- ules—“vector spaces” whose coefficient are drawn from a ring—or as additive abelian groups. 126 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY P2P3P4 Figure 4.5: The oriented triangle P2P3P4has boundary P3P4+P4P2+P2P3. The Boundary Operator We now introduce a linear map ∂p:Cp→Cp−1, called the boundary operator . Its action on a p-simplex is ∂pPi1Pi2···Pip+1=p+1/summationdisplay j=1(−1)j+1Pi1.../hatwidePij...Pip+1, (4.23) where the “hat” indicates that Pijis to be omitted. The resulting ( p−1)- chain is called the boundary of the simplex. For example ∂2(P2P3P4) =P3P4−P2P4+P2P3, =P3P4+P4P2+P2P3. (4.24) The boundary of a line segment is the difference of its endpoin ts ∂1(P1P2) =P2−P1. (4.25) Finally, for any point, ∂Pi= 0. (4.26) Because∂is defined to be a linear map, when it is applied to a p-chain c=a1s1+a2s2+···+ansn, where the siarep-simplices, we have ∂pc= a1∂ps1+a2∂ps2+···+an∂psn. When we take the “ ∂” of a chain of compatibly oriented simplices that to- gether make up some region, the internal boundaries cancel i n pairs, and the “boundary” of the chain really is the oriented geometric boundary of the region. For example in figure 4.6 we find that ∂(P1P5P2+P2P5P4+P3P4P5+P1P3P5) =P1P3+P3P4+P4P2+P2P1,(4.27) 4.3. HOMOLOGY 127 P42P P1 P3P5 Figure 4.6: Compatibly oriented simplices. which is the counter-clockwise directed boundary of the squ are. For each of the examples we find that ∂p−1∂ps= 0. From the definition (4.23) we can easily establish that this identity holds for a nyp-simplexs. As chains are sums of simplices and ∂pis linear, it remains true for any c∈Cp. Thus∂p−1∂p= 0. We will usually abbreviate this statement as ∂2= 0. Cycles, Boundaries and Homology Achain complex is a doubly infinite sequence of spaces (these can be vector spaces, modules, abelian groups, or many other mathematica l objects) such as...,C −2,C−1,C0,C1,C2..., together with structure-preserving maps ...∂p+1→Cp∂p→Cp−1∂p−1→Cp−2∂p−1→..., (4.28) with the property that ∂p−1∂p= 0. The finite sequence of Cp’s we constructed from our simplicial complex is an example of a chain complex w hereCpis zero-dimensional for p <0 orp > d . Chain complexes are a useful tool in mathematics, and the ideas we explain in this section have ma ny applications. Given any chain complex we can define two important linear sub spaces of each of the Cp’s. The first is the space Zpofp-cycles . This consists of thosez∈Cpsuch that∂pz= 0. The second is the space Bpofp-boundaries , and consists of those b∈Cpsuch thatb=∂p+1cfor somec∈Cp+1. Because ∂2= 0, the boundaries Bpconstitute a subspace of Zp. From these spaces we form the quotient space Hp=Zp/Bp, consisting of equivalence classes of p-cycles, where we deem z1andz2to be equivalent, or homologous , if they differ by a boundary: z2=z1+∂c. We will write the equivalence class of cycles homologous zito as [zi]. The space Hp, or more accurately, Hp(R), is called thep-th (simplicial) homology space of the chain complex. It becomes thep-th homology group ifRis replaced by the integers. 128 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY We can construct these homology spaces for any chain complex . When the chain complex is derived from a simplicial complex decom position of a manifoldMa remarkable thing happens. The spaces Cp,Zp, andBp, all depend on the details of how the manifold Mhas been dissected to form the simplicial complex S. The homology space Hp, however, is independent the dissection. This is neither obvious nor easy to prove. We will rely on examples to make it plausible. Granted this independence, w e will write Hp(M), orHp(M,R), so as to make it clear that Hpis a property of M. The dimensionbpofHp(M) is called the p-thBetti number of the manifold: bpdef= dimHp(M). (4.29) Example: The Two-Sphere. For the tetrahedral dissection of the two-sphere, any vertex is Pihomologous to any other, as Pi−Pj=∂(PjPi) and all PjPibelong toC2. Furthermore, ∂Pi= 0, soH0(S2) is one dimensional. In general, the dimension of H0(M) is the number of disconnected pieces making up M. We will write H0(S2) =R, regarding Ras the archetype of a one-dimensional vector space. Now let us consider H1(S2). We first find the space of 1-cycles Z1. An element ofC1will be inZ1only if each vertex that is the begining of an edge is also the end of an edge, and that these edges have the same co efficient. Thus z1=P2P3+P3P4+P4P2 is a cycle, as is z2=P1P4+P4P2+P2P1. These are both boundaries of faces of the tetrahedron. It sho uld be fairly easy to convince yourself that Z1is the space of linear combinations of these together with boundaries of the other faces z3=P1P4+P4P3+P3P1, z4=P1P3+P3P2+P2P1. Any three of these are linearly independent, and so Z1is three dimensional. Because all of the cycles are boundaries, every element of Z1is homologous to0, and soH1(S2) ={0}. We also see that H2(S2) =R. Here the basis element is P2P3P4−P1P3P4+P1P2P4−P1P2P3 (4.30) 4.3. HOMOLOGY 129 which is the 2-chain corresponding to the entire surface of t he sphere. It would be the boundary of the solid tedrahedron, but does not c ount as a boundary as the interior of the tetrahedron is not part of the simplicial complex. Example: The Torus. Consider the 2-torus T2.We will see that H0(T2) =R, H1(T2) =R2≡R⊕R, andH2(T2) =R. A natural basis for the two- dimensional H1(T2) consists of the 1-cycles α,βportrayed in figure 4.7. αβ Figure 4.7: A basis of 1-cycles on the 2-torus. The cycleγthat, in figure 4.2, winds once around the torus is homologous toα+β. In terms of the second triangulation of the torus (figure 4.3 ) we would have α=P1P2+P2P3+P3P1 β=P1P7+P7P4+P4P1 (4.31) and γ=P1P8+P8P6+P6P1 =α+β+∂(P1P8P2+P8P9P2+P2P9P3+···). (4.32) Example: The Projective Plane. The projective plane RP2can be regarded as a rectangle with diametrically opposite points identifie d. Suppose we decompose RP2into eight triangles, as in figure 4.8. 130 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY P1P1 PP P2P2 3P4P43 P5 Figure 4.8: A triangulation of the projective plane. Consider the “entire surface” σ=P1P2P5+P1P5P4+···∈C2(RP2), (4.33) consisting of the sum of all eight 2-simplices with the orien tation indicated in the figure. Let α=P1P2+P2P3andβ=P1P4+P4P3be the sides of the rectangle running along the bottom horizontal and left vert ical sides of the figure, respectively. In each case they run from P1toP3. Then ∂(σ) =P1P2+P2P3+P3P4+P4P1+P1P2+P2P3+P3P4+P1P2 = 2(α−β)/negationslash= 0. (4.34) Although RP2has no actual edge that we can fall off, from the homological viewpoint it does have a boundary! This represents the confli ct between local orientation of each of the 2-simplices and the global non-or ientability of RP2. The surface σofRP2is not a two-cycle, therefore. Indeed Z2(RP2), and a fortioriH2(RP2), contain only the zero vector. The only one-cycle is α−β which runs from P1toP1viaP2,P3andP4, but (4.34) shows that this is the boundary of1 2σ. ThusH2(RP2,R) ={0}andH1(RP2,R) ={0}, while H0(RP2,R) =R. We can now see the advantage of restricting ourselves to inte ger coeffi- cients. When we are not allowed fractions, the cycle γ= (α−β) is no longer a boundary, although 2( α−β) is the boundary of σ. Thus, using the symbol Z2to denote the additive group of the integers modulo two, we can write H1(RP2,Z) =Z2. This homology space is a set with only two members {0γ,1γ}. The finite group H1(RP2,Z) =Z2is said to be the torsion part of the homology — a confusing terminology because this torsi on has nothing to do with the torsion tensor of Riemannian geometry. 4.3. HOMOLOGY 131 We introduced real-number homology first, because the theor y of vector spaces is simpler than that of modules, and more familiar to p hysicists. The torsion is, however, invisible to the real-number homology . We were therefore buying a simplification at the expense of throwing away infor mation. The Euler Character The sum χ(M)def=d/summationdisplay p=0(−1)pdimHp(M,R) (4.35) is called the Euler character of the manifold M. For example, the 2-sphere hasχ(S2) = 2, the projective plane has χ(RP2) = 1, and the n-torus has χ(Tn) = 0. This number is manifestly a topological invariant beca use the individual dim Hp(M,R) are. We will show that that the Euler character is also equal to V−E+F−··· whereVis the number of vertices, Eis the number of edges and Fis the number of faces in the simplicial dissection. The dots are for higher dimensional spaces, where the alternati ng sum continues with (−1)ptimes the number of p-simplices. In other words, we are claiming that χ(M) =d/summationdisplay p=0(−1)pdimCp(M). (4.36) It is not so obvious that this new sum is a topological invaria nt. The indi- vidual dimensions of the spaces of p-chains depend on the details of how we dissectMinto simplices. If our claim is to be correct, the dependence must somehow drop out when we take the alternating sum. A useful tool for working with alternating sums of vector-sp ace dimen- sions is provided by the notion of an exact sequence . We say that a set of vector spaces Vpwith maps fp:Vp→Vp+1is an exact sequence if Ker (fp) = Im (fp−1). For example, if all cycles were boundaries then the set of spaces Cpwith the maps ∂ptaking us from CptoCp−1would consi- tute an exact sequence—albeit with pdecreasing rather than increasing, but this is irrelevent. When the homology is non-zero, however, we only have Im (fp−1)⊂Ker (fp), and the number dim Hp= dim (Ker fp)−dim (Imfp−1) provides a measure of how far this set inclusion falls short o f being an equal- ity. 132 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY Suppose that {0}f0−→V1f1−→V2f2−→...fn−1−→Vnfn−→{0} (4.37) is a finite-length exact sequence. Here, {0}is the vector space containing only the zero vector. Being linear, f0maps0to0. Alsofnmaps everything inVnto0. Since this last map takes everything to zero, and what is map ped to zero is the image of the penultimate map, we have Vn= Imfn−1. Similarly, the fact that Ker f1= Imf0={0}shows that Im f1⊆V2is an isomorphic image ofV1. This situation is represented pictorially in figure 4.9. }{V1V2V3 V4V5 fIm Imf ImfImf}{f0f f f f4 f50 0 0 02 1 3 401 2 3 0 0 0 0 Figure 4.9: A schematic representation of an exact sequence. Now the range-nullspace theorem tells us that dimVp= dim (Im fp) + dim (Ker fp) = dim (Im fp) + dim (Im fp−1). (4.38) When we take the alternating sum of the dimensions, and use di m (Imf0) = 0 and dim (Im fn) = 0, we find that the sum telescopes to give n/summationdisplay p=0(−1)pdimVp= 0. (4.39) The vanishing of this alternating sum is one of the principal properties of an exact sequence. Now, for our sequence of spaces Cpwith the maps ∂p:Cp→Cp−1, we have dim (Ker∂p) = dim (Im ∂p+1) + dimHp. Using this and the range-nullspace 4.3. HOMOLOGY 133 theorem in the same manner as above, shows that d/summationdisplay p=0(−1)pdimCp(M) =d/summationdisplay p=0(−1)pdimHp(M). (4.40) This confirms our claim. Exercise 4.1 : Count the number of vertices, edges, and faces in the triang u- lation we used to compute the homology groups of the real proj ective plane RP2. Verify that V−E+F= 1, and that this is the same number that we get by evaluating χ(RP2) = dimH0(RP2,R)−dimH1(RP2,R) + dimH2(RP2,R). Exercise 4.2 : Show that the sequence {0}→Vφ→W→{0} of vector spaces being exact means that the map φ:V→Wis one-to-one and onto, and hence an isomorphism V∼=W. Exercise 4.3 : Show that a short exact sequence {0}→Ai→Bπ→C→{0} of vector spaces is just a sophisticated way of asserting tha tC∼=B/A. More precisely, show that the map iis injective (one-to-one), so Acan be considered to be a subspace of B. Then show that the map πis surjective (onto), and can be regarded as projecting Bonto the equivalence classes B/A. Exercise 4.4 : Letα:A→Bbe a linear map. Show that {0}→Kerαi→Aα→Bπ→Cokerα→{0} is an exact sequence. (Recall that Coker α≡B/Imα.) 4.3.2 Relative homology Mathematicians have invented powerful tools for computing homology. In this section we introduce one of them: the exact sequence of a pair . We 134 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY describe this tool in detail because a homotopy analogue of t his exact se- quence is used in physics to classify defects such as disloca tions, vortices and monopoles. Homotopy theory is however harder and requires m ore technical apparatus than homology, so the ideas are easier to explain h ere. We have seen that it is useful to think of complicated manifol ds as being assembled out of simpler ones. We constructed the torus, for example, by gluing together edges of a rectangle. Another construction technique involves shrinking parts of a manifold to a point. Think, for example, of the unit 2- disc as a being circle of cloth with a drawstring sewn into its boundary. Now pull the string tight to form a spherical bag. The continuous functions on the resulting 2-sphere are those continuous functions on th e disc that took the same value at all points on its boundary. Recall that we us ed this idea in 3.4.2, where we claimed that those spin textures in R2that point in a fixed direction at infinity can be thought of as spin textures on the 2-sphere. We now extend this shrinking trick to homology. Suppose that we have a chain complex consisting of spaces Cpand bound- ary operations ∂p. We wiill denote this chain complex by ( C,∂). Another set of of spaces and boundary operations ( C/prime,∂/prime) is asubcomplex of (C,∂) if eachC/prime p⊆Cpand∂/prime p(c) =∂p(c) for eachc∈C/prime p. This situation arises if we have a simplical complex Sand a some subset S/primethat is itself a simplicial complex, and take C/prime p=Cp(S/prime) Since each C/prime pis subspace of Cpwe can form the quotient spaces Cp/C/prime p and make them into a chain complex by defining, for c+C/prime p∈Cp/C/prime p, ∂p(c+C/prime p) =∂pc+C/prime p−1. (4.41) It easy to see that this operation is well defined ( i.e.it gives the same output independent of the choice of representative in the equivale nce classc+C/prime p), that∂p:Cp→Cp−1is a linear map, and that ∂p−1∂p= 0. We have constructed a new chain complex ( C/C/prime,∂). We can therefore form its ho- mology spaces in the usual way. The resulting vector space, o r abelian group, Hp(C/C/prime) is thep-threlative homology group of CmoduloC/prime. WhenC/primeand Carise from simplicial complexes S/prime⊆S, these spaces are what remains of the homology of Safter every chain in S/primehas been shrunk to a point. In this case, it is customary to write Hp(S,S/prime) instead of Hp(C/C/prime), and simi- larly write the chain, cycle and boundary spaces as Cp(S,S/prime),Zp(S,S/prime) and Bp(S,S/prime) respectively. Example: Constructing the two-sphere S2from the two-ball (or disc) B2. We regard B2to be the triangular simplex P1P2P3, and its boundary, the 4.3. HOMOLOGY 135 one-sphere or circle S1, to be the simplicial complex containing the points P1, P2,P3and the sides P1P2,P2P3,P3P1, but not the interior of the triangle. We wish to contract this boundary complex to a point, and form the relative chain complexes and their homology spaces. Of the spaces we q uotient by, C0(S1) is spanned by the points P1,P2,P3, the 1-chain space C1(S1) is spanned by the sides P1P2,P2P3,P3P1, whileC2(S1) ={0}. The space of relative chains C2(B1,S1) consists of multiples of P1P2P3+C2(S1), and the boundary ∂2/parenleftBig P1P2P3+C2(S1)/parenrightBig = (P2P3+P3P1+P1P2) +C1(S1) (4.42) is equivalent to zero because P2P3+P3P1+P1P2∈C1(S1). ThusP1P2P3+ C2(S1) is a non-bounding cycle and spans H2(B2,S1), which is therefore one dimensional. This space is isomorphic to the one-dimens ionalH2(S2). SimilarlyH1(B2,S1) is zero dimensional, and so isomorphic to H1(S2). This is because all chains in C1(B2,S1) are inC1(S1) and therefore equivalent to zero. A peculiarity, however, is that H0(B2,S1) isnotisomorphic to H0(S2) = R. Instead, we find that H0(B2,S1) ={0}because all the points are equiva- lent to zero. This vanishing is characteristic of the zeroth relative homology spaceH0(S,S/prime) for the simplicial triangulation of any connected manifol d. It occurs because Sbeing connected means that any point PinScan be reached by walking along edges from any other point, in parti cular from a pointP/primeinS/prime. This makes Phomologous to P/prime, and so equivalent to to zero inH0(S,S/prime). Exact homology sequence of a pair Homological algebra is full of miracles. Here we describe on e of them. From the ingredients we have at hand, we can construct a semi-infin ite sequence of spaces and linear maps between them ···∂∗p+1−→Hp(S/prime)i∗p−→Hp(S)π∗p−→Hp(S,S/prime)∂∗p−→ Hp−1(S/prime)i∗p−1−→Hp−1(S)π∗p−1−→Hp−1(S,S/prime)∂∗p−1−→ ... ∂∗1−→H0(S/prime)i∗0−→H0(S)π∗0−→H0(S,S/prime)∂∗0−→{0}.(4.43) 136 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY The mapsi∗pandπ∗pare induced by the natural injection ip:Cp(S/prime)→Cp(S) and projection πp:Cp(S)→Cp(S)/Cp(S/prime). It is only necessary to check that πp−1∂p=∂pπp, ip−1∂p=∂pip, (4.44) to see that they are compatible with the passage from the chai n spaces to the homology spaces. More discussion is required of the connection map ∂∗p that takes us from one row to the next in the displayed form of ( 4.43). Leth∈Hp(S,S/prime), thenh=z+Bp(S,S/prime) for some cycle z∈Z(S,S/prime), and in turnz=c+Cp(S/prime) for somec∈Cp(S). (So twochoices of representative of equivalence class are being made here.) Now ∂pz= 0 which means that ∂pc∈Cp−1(S/prime). This fact, when combined with ∂p−1∂p= 0, tells us that ∂pc∈Zp−1(S/prime). We now set ∂∗p(h) =∂pc+Bp−1(S/prime). (4.45) This sounds rather involved, but let’s say it again in words: an element of Hp(S,S/prime) is a relative p-cycle moduloS/prime. This means that its boundary is not necessarily zero, but may be a non-zero element of Cp−1(S/prime). Since this element is the boundary of something its own boundary vanish es, so it is (p−1)-cycle in Cp−1(S/prime) and hence a representative of a homology class in Hp−1(S/prime). This homology class is the output of the ∂∗pmap. The miracle is that the sequence of maps (4.43) is exact. It is an example of a standard homological algebra construction of a long exact sequence out of a family of short exact sequences, in this case out the sequ ences {0}→Cp(S/prime)→Cp(S)→Cp(S,S/prime)→{0}. (4.46) Proving that the long sequence is exact is straightforward. All one must do is check each map to see that it has the properties required. T his exercise in diagram chasing is left to the reader. This long exact sequence is called the exact homology sequence of a pair . If we know that certain homology spaces are zero dimensional , it provides a powerful tool for computing other spaces in the sequence. As an illustration, consider the sequence of the pair Bn+1andSnforn>0: ···i∗p−→Hp(Bn+1)/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright ={0}π∗p−→Hp(Bn+1,Sn)∂∗p−→Hp−1(Sn) 4.3. HOMOLOGY 137 i∗p−1−→Hp−1(Bn+1)/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright ={0}π∗p−1−→Hp−1(Bn+1,Sn)∂∗p−1−→Hp−2(Sn) ... i∗1−→H1(Bn+1)/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright ={0}π∗1−→H1(Bn+1,Sn)∂∗1−→H0(Sn)/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright =R i∗0−→H0(Bn+1)/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright =Rπ∗0−→H0(Bn+1,Sn)∂∗0−→{0}. (4.47) We have inserted here the easily established data that Hp(Bn+1) ={0}for p>0 (which is a consequence of the ( n+1)-ball being a contractible space), and thatH0(Bn+1) andH0(Sn) are one dimensional because they consist of a single connected component. We read off, from the {0}→A→B→{0} exact subsequences, the isomorphisms Hp(Bn+1,Sn)∼=Hp−1(Sn), p> 1, (4.48) and from the exact sequence {0}→H1(Bn+1,S1)→R→R→H0(Bn+1,Sn)→{0} (4.49) thatH1(Bn+1,Sn) ={0}=H0(Bn+1,Sn). The first of these equalities holds becauseH1(Bn+1,Sn) is the kernel of the isomorphism R→R, and the second because H0(Bn+1,Sn) is the range of a surjective null map. In the casen= 0, we have to modify our last conclusion because H0(S0) = R⊕Ris two dimensional. (Remember that H0(M) counts the number of disconnected components of M, and the zero-sphere S0consists of the two disconnected points P1,P2lying in the boundary of the interval B1=P1P2.) As a consequence, the last five maps become {0}→H1(B1,S0)→R⊕R→R→H0(B1,S0)→{0}. (4.50) This tells us that H1(B1,S0) =RandH0(B1,S0) ={0}. Exact homotopy sequence of a pair We have met the homotopy groups πn(M) in section 3.4.4. As we saw there, homotopy groups can be used to classify defects or solitons i n physical sys- tems in which some field takes values in the manifold M. When the system 138 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY has undergone spontaneous symmetry breaking from a larger s ymmetryG to a subgroup H, the relevant manifold is the coset G/H. The group πn(G) can be taken to be the set of continuous maps of an n-dimensional cube into G, with the surface of the cube mapping to the identity element e∈G. We similarly define the relative homotopy group πn(G,H) ofGmoduloHto be the set of continuous maps of the cube into G, with all-but-one face of the cube mapping to e, but with the remaining face mapping to the subgroup H. It can then be shown that πn(G/H)/similarequalπn(G,H) (the hard part is to show that any continuous map into G/H can be represented as the projection of some continuous map into G). The short exact sequence {e}→Hi→Gπ→G/H→{e} (4.51) of group homomorphisms (where {e}is the group consisting only of the identity element) then gives rise to the long exact sequence ···→πn(H)→πn(G)→πn(G,H)→πn−1(H)→··· (4.52) The derivation and utility of this exact sequence is very wel l described in the review article by Mermin cited in section 3.4.4. We have ther efore contented ourselves with simply displaying the result so that the read er can see the similarity between the homology theorem and its homotopy-t heory analogue. 4.4 De Rham’s Theorem We still have not related homology to cohomology. The link is provided by integration. The integral provides a natural pairing of a p-chaincand ap-formω: if c=a1s1+a2s2+···+ansn, where the siare simplices, we set (c,ω) =/summationdisplay iai/integraldisplay siω. (4.53) The perhaps mysterious notion of “adding” geometric simpli ces is thus given a concrete interpretation in terms of adding real numbers. Stokes’ theorem now reads (∂c,ω) = (c,dω), (4.54) 4.4. DE RHAM’S THEOREM 139 suggesting that dand∂should be regarded as adjoints of each other. From this observation follows the key fact that the pairing betwe en chains and forms descends to a pairing between homology classes and coh omology classes. In other words, (z+∂c,ω+dχ) = (z,ω), (4.55) so it does not matter which representative of the equivalenc e classes we take when we compute the integral. Let us see why this is so: Supposez∈Zpandω2=ω1+dη. Then (z,ω2) =/integraldisplay zω2=/integraldisplay zω1+/integraldisplay zdη =/integraldisplay zω1+/integraldisplay ∂zη =/integraldisplay zω1 = (z,ω1) (4.56) because∂z= 0. Thus, all elements of the cohomology class of ωreturn the same answer when integrated over a cycle. Similarly, if ω∈Zpandc2=c1+∂athen (c2,ω) =/integraldisplay c1ω+/integraldisplay ∂aω =/integraldisplay c1ω+/integraldisplay adω =/integraldisplay c1ω = (c1,ω), sincedω= 0. All this means that we can consider the equivalence classes o f closed forms composing Hp DR(M) to be elements of ( Hp(M))∗, the dual space of Hp(M) — hence the “co” in cohomology. The existence of the pairing d oes not automatically mean that Hp DRisthe dual space to Hp(M), however, because there might be elements of the dual space that are not in Hp DR, and there might be distinct elements of Hp DRthat give identical answers when integrated over any cycle, and so correspond to the same element in ( Hp(M))∗. This 140 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY does not happen, however, when the manifold is compact : De Rham showed that, for compact manifolds, ( Hp(M,R))∗=Hp DR(M,R). We will not try to prove this, but be satisfied with some examples. The statement ( Hp(M))∗=Hp DR(M) neatly summarizes de Rham’s re- sults, but, in practice, the more explicit statements below are more useful. Theorem: (de Rham) Suppose that Mis a compact manifold. 1) A closed p-formωis exact if and only if /integraldisplay ziω= 0 (4.57) for all cycles zi∈Zp. It suffices to check this for one representative of each homology class. 2) Ifzi∈Zp,i= 1,...,dimHp, is a basis for the p-th homology space, andαia set of numbers, one for each zi, then there exists a closed p-formωsuch that /integraldisplay ziω=αi. (4.58) Ifωiconstitute a basis of the vector space Hp(M) then the matrix of numbers Ωij= (zi,ωj) =/integraldisplay ziωj(4.59) is called the period matrix , and the Ω ijthemselves are the periods . Example:H1(T2) =R⊕Ris two-dimensional. Since a finite-dimensional vector space and its dual have the same dimension, de Rham tel ls us that H1 DR(T2) is also two-dimensional. If we take as coordinates on T2the angles θandφ, then the basis elements, or generators , of the cohomology spaces are the forms “ dθ” and “dφ”. We have inserted the quotes to stress that these expressions are not the dof a function. The angles θandφarenotfunctions on the torus, since they are not single-valued. The homology basis 1-cycles can be taken as zθrunning from θ= 0 toθ= 2πalongφ=π, andzφrunning fromφ= 0 toφ= 2πalongθ=π. Clearly,ω=αθdθ/2π+αφdφ/2πreturns/integraltext zθω=αθand/integraltext zφω=αφfor anyαθ,απ, so{dθ/2π,dφ/ 2π}and{zθ,zφ}are dual bases. Example: We have earlier computed H2(RP2,R) ={0}andH1(RP2,R) = {0}. De Rham therefore tells us that H2(RP2,R) ={0}andH1(RP2,R) = {0}. From this we deduce that all closed one- and two-forms on the projective plane RP2are exact. 4.4. DE RHAM’S THEOREM 141 Example : As an illustration of de Rham part 1), observe that it is easy to show that a closed one-forms φcan be written as df, provided that/integraltext ziφ= 0 for all cycles. We simply define f=/integraltextx x0φ, and observe that the proviso ensures that fis not multivalued. Example : A more subtle problem is to show that, given a two-form ωonS2, with/integraltext S2ω= 0, then there is a globally defined χsuch thatω=dχ. We begin by covering S2by two open sets D+andD−which have the form of caps such that D+includes all of S2except for a neighbourhood of the south pole, while D−includes everything except a neighbourhood of the north pol e, and the intersection, D+∩D−, has the topology of an annulus, or cingulum , encircling the equator. D D+ _Γ Figure 4.10: A covering the sphere by two contractable caps. Since bothD+andD−are contractable, there are one-forms χ+andχ−such thatω=dχ+inD+andω=dχ−inD−. Thus, d(χ+−χ−) = 0,inD+∩D−. (4.60) Dividing the sphere into two disjoint sets with a common (but oppositely oriented) boundary Γ ∈D+∩D−we have 0 =/integraldisplay S2ω=/contintegraldisplay Γ(χ+−χ−), (4.61) and this is true for any such curve Γ. Thus, by the previous exa mple, φ≡(χ+−χ−) =df (4.62) for some smooth function fdefined inD+∩D−. We now introduce a partition of unity subordinate to the cover of S2byD+andD−. This partition is a 142 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY pair of non-negative smooth functions, ρ±, such that ρ+is non-zero only in D+,ρ−is non-zero only in D−, andρ++ρ−= 1. Now f=ρ+f−(−ρ−)f, (4.63) andf−=ρ+fis a function defined everywhere on D−. Similarly f+= (−ρ−)fis a function on D+. Notice the interchange of ±labels! This is not a mistake. The function fis not defined outside D+∩D−, but we can define ρ−feverywhere on D+becausefgets multiplied by zero wherever we have no specific value to assign to it. We now observe that χ++df+=χ−+df−,inD+∩D−. (4.64) Thusω=dχ,whereχis defined everywhere by the rule χ=/braceleftbiggχ++df+,inD+, χ−+df−,inD−.(4.65) It does not matter which definition we take in the cingular reg ionD+∩D−, because the two definitions coincide there. The methods of this example, a special case of the Mayer-Vietoris prin- ciple, can be extended to give a proof of de Rham’s claims. 4.5 Poincar´ e Duality De Rham’s theorem does not require that our manifold Mbe orientable. Our next results do, however, require orientablity. We therefo re assume through- out this section that Mis a compact, orientable, D-dimensional manifold. We begin with the observation that if the forms ω1andω2are closed then so isω1∧ω2. Furthermore if one or both of ω1,ω2is exact then the product ω1∧ω2is also exact. It follows that the cohomology class [ ω1∧ω2] ofω1∧ω2 depends only on the cohomology classes [ ω1] and [ω2]. The wedge product thus induces a map Hp(M,R)×Hq(M,R)∧→Hp+q(M,R), (4.66) which is called the “cup product” of the cohomology classes. It is written as [ω1∧ω2] = [ω1]∪[ω2], (4.67) 4.5. POINCAR ´E DUALITY 143 and gives the cohomology the structure of a graded-commutat ive ring, de- noted byH•(M,R) More significant for us than the ring structure is that, given ω∈HD(M,R), we can obtain a real number by forming/integraltext Mω(This is the point at which we need orientability. We only know how to integrate over ori entable chains, and so cannot even define/integraltext MωwhenMis not orientable.) and can com- bine this integral with the cup product to make any cohomolog y class [f]∈ HD−p(M,R) into an element Fof (Hp(M,R))∗. We do this by setting F([g]) =/integraldisplay Mf∧g (4.68) for each [g]∈Hp(M,R). Furthermore, it is possible to show that we can getanyelementFof (Hp(M,R))∗in this way, and the corresponding [ f] is unique . But de Rham has already given us a way of identifying the elem ents of (Hp(M,R))∗with the cycles in Hp(M,R)! There is, therefore, a 1-1 onto map Hp(M,R)↔HD−p(M,R). (4.69) In particular the dimensions of these two spaces must coinci de bp(M) =bD−p(M). (4.70) This equality of Betti numbers is called Poincar´ e duality . Poincar´ e originally conceived of it geometrically. His idea was to construct fro m each simplicial triangulation SofMa new “dual” triangulation S/prime, where, in two dimensions for example, we place a new vertex at the centre of each triang le, and join the vertices by lines through each side of the old triangles to ma ke new cells — each new cell containing one of the old vertices. If we are luc ky, this process will have the effect of replacing each p-simplex by a ( D−p)-simplex, and so set up a map between Cp(S) andCD−p(S/prime) that turns the homolgy “upside down.” The new cells are not always simplices, however, and i t is hard to make this construction systematic. Poincar´ e’s original r ecipe was flawed. Our present approach to Poincar´ e’s result is asserting tha t for each basis p-cycle class [ zp i] there is a unique (up to cohomology) ( D−p)-formωD−p i such that /integraldisplay zp if=/integraldisplay MωD−p i∧f. (4.71) We can construct this ωD−p i“physically” by taking a representative cycle zp i in the homology class [ zp i] and thinking of it as a surface with a conserved 144 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY unit (d−p)-form current flowing in its vicinity. An example would be th e two-form topological current running along the one-dimens ional worldline of a Skyrmion. (See the discussion surrounding equation (3.63 ).) TheωD−p i form a basis for HD−p(M,R). We can therefore expand f∼fiωD−p i,and similarly for the closed p-formg, to obtain /integraldisplay Mg∧f=figjI(i,j) (4.72) where the matrix I(i,j)≡I(zp i,zD−p j) =/integraldisplay MωD−p i∧ωp j (4.73) is called the intersection form . From the definition we have I(i,j) = (−1)p(D−p)I(j,i). (4.74) Less obvious is that I(i,j) is an integer that reports the number of times (counted with orientation) that the cycles zp iandzD−p jintersect. This latter fact can be understood from our construction of the ωp ias unit currents localized near the zD−p icycles. The integrand in (4.73) is non-zero only in the neighbourhood of the intersections of zp iwithzD−p j, and at each intersection constitutes a D-form that integrates up to give ±1. +1 +1 −1 +1 α αβ β Figure 4.11: The intersection of two cycles: I(α,β) = 1 = 1−1 + 1. This claim is illustrated in the left-hand part of figure 4.11 , which shows a region surrounding the intersection of the αandβone-cycles on the 2-torus. The co-ordinate system has been chosen so that the αcycle runs along the 4.5. POINCAR ´E DUALITY 145 xaxis and the βcycle along then yaxis. Each cycle is surrounded by the narrow shaded regions −w < y < w and−w < x < w , respectively. To construct suitable forms ωαandωβwe select a smooth function f(x) that vanishes for|x|≥wand such that/integraltext fdx= 1. In the local chart we can then set ωα=f(y)dy, ωβ=−f(x)dx, both these forms being closed. The intersection number is gi ven by the integral I(α,β) =/integraldisplay ωα∧ωβ=/integraldisplay/integraldisplay f(x)f(y)dxdy= 1. (4.75) The right-hand part of figure 4.11 illustrates why this inter section number depends only on the homology classes of the two one-cycles, a nd not on their particular instantiation as curves. We can more conveniently re-express (4.72) terms of the periods of the forms fi≡/integraldisplay zp if=I(i,k)fk, gj≡/integraldisplay zD−p jg=I(j,l)gl, (4.76) as /integraldisplay Mf∧g=/summationdisplay i,jK(i,j)/integraldisplay zp if/integraldisplay zD−p jg, (4.77) where K(i,j) =I−1(i,k)I−1(j,l)I(k,l) =I−1(j,i) (4.78) is the transpose of the inverse of the intersection-form mat rix. The decom- position (4.77) of the integral of the product of a pair of clo sed forms into a bilinear form in their periods is one of the two principal re sults of this section, the other being (4.70). In simple cases we can obtain the decomposition (4.77) by mor e direct methods. Suppose, for example, that we label the cycles gene rating the homology group H1(T2) of the 2-torus as αandβ, and that aandbare closed (da=db= 0), but not necessarily exact, one-forms. We will show that /integraldisplay T2a∧b=/integraldisplay αa/integraldisplay βb−/integraldisplay αb/integraldisplay βa. (4.79) 146 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY To do this, we cut the torus along the cycles αandβand open it out into a rectangle with sides of length LxandLy. The cycles αandβwill form the sides of the rectangle and we will take them as lying paral lel to thex andyaxes, respectively. Functions on the torus now become functions on therectangle . Not all functions on the rectangle descend from functions o n the torus, however. Only those functions that satisfy the pe riodic bound- ary conditions f(0,y) =f(Lx,y) andf(x,0) =f(x,Ly) can be considered (mathematicians would say “can be lifted”) to be functions on the torus. 2T ααα ββ β Figure 4.12: Cut-open torus Since the rectangle (but not the torus) is retractable, we ca n writea=df wherefis a function on the rectangle — but not necessarily a functio n on the torus, i.e.,fwill not, in general, be periodic. Since a∧b=d(fb), we can now use Stokes’ theorem to evaluate /integraldisplay T2a∧b=/integraldisplay T2d(fb) =/integraldisplay ∂T2fb. (4.80) The two integrals on the two vertical sides of the rectangle c an be combined to a single integral over the points of the one-cycle β: /integraldisplay verticalfb=/integraldisplay β[f(Lx,y)−f(0,y)]b. (4.81) We now observe that [ f(Lx,y)−f(0,y)] is a constant, and so can be taken out of the integral. It is a constant because all paths from th e point (0,y) to (Lx,y) are homologous to the one-cycle α, so the difference f(Lx,y)−f(0,y) is equal to/integraltext αa. Thus /integraldisplay β[f(Lx,y)−f(0,y)]b=/integraldisplay αa/integraldisplay βb. (4.82) 4.6. CHARACTERISTIC CLASSES 147 Similarly, the contributions of the two horizontal sides is /integraldisplay α[f(x,0)−f((x,Ly)]b=−/integraldisplay βa/integraldisplay αb. (4.83) On putting the contributions of both pairs of sides together , the claimed result follows. 4.6 Characteristic Classes A supply of elements of H2m(M,R) andH2m(M,Z) is provided by the charac- teristic classes associated with connections on vector bundles over the man- ifoldM. Recall that connections appear in covariant derivatives ∇µ≡∂µ+Aµ, (4.84) and are to be thought of as matrix-valued one-forms A=Aµdxµ. In the quantum mechanics of charged particles the covariant deriv ative that appears in the Schr¨ odinger equation is ∇µ=∂ ∂xµ−ieAMaxwell µ. (4.85) Hereeis the charge of the particle on whose wavefunction the deriv ative acts, andAMaxwell µ is the usual electromagnetic vector potential. The matrix- valued connection one-form is therefore A=−ieAMaxwell µdxµ. (4.86) In this case the matrix is one-by-one. In a non-abelian gauge theory with gauge group Gthe connection becomes A=iˆλaAa µdxµ(4.87) Theˆλaare hermitian matrices that have commutation relations [ ˆλa,ˆλb] = ifc abˆλc, where the fc abare the structure constants of the Lie algebra of G. The ˆλatherefore form a representation of the Lie algebra, and this representation plays the role of the “charge” of the non-abelian gauge parti cle. 148 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY For covariant derivatives acting on a tangent vector field faeaon a Rie- mannn-manifold, where the eaare an orthonormal vielbein frame, we have A=ωabµdxµ, (4.88) where, for each µ, the coefficients ωabµ=−ωbaµcan be thought of as the entries in a skew symmetric n-by-nmatrix. These matrices are elements of the Lie algebra o(n) of O(n). In all these cases we define the curvature two-form to be F=dA+A2, where a combined matrix and wedge product is to be understood inA2. In exercises 2.19 and 2.20 you used the Bianchi identity to sh ow that the gauge-invariant 2 n-forms tr(Fn) were closed. The integrals of these forms over cycles provide numbers that are topological invariant s of the bundle. For example, in four-dimensional QCD, the integral c2=−1 8π2/integraldisplay Ωtr (F2), (4.89) over a compactified four-dimensional manifold Ω is an intege r that a math- ematician would call the second Chern number of the non-abel ian gauge bundle, and that a physicist would call the instanton number of the gauge field configuration. In this section we will show that the integrals of such charac teristic classes are indeed topological invariants. We also explain somethi ng of what these invariants are measuring, and illustrate why, when suitabl y normalized, cer- tain of them are integer valued. 4.6.1 Topological invariance Suppose that we have been given a connection Aand slightly deform it A→A+δA, then δF=d(δA) +δAA+AδA. (4.90) Using the Bianchi identity dF=FA−AF, we find that δtr(Fn) =ntr(δFFn−1) =ntr(d(δA)Fn−1) +ntr(δAAFn−1) +ntr(AδAFn−1) =ntr(d(δA)Fn−1) +ntr(δAAFn−1)−ntr(δAFn−1A) =d/braceleftbig ntr(δAFn−1)/bracerightbig . (4.91) 4.6. CHARACTERISTIC CLASSES 149 The last line of (4.91) is equal to the penultimate line becau se all but the first and last terms arising from the dF’s ind{tr(δAFn−1)}cancel in pairs. A globally defined change in Atherefore changes tr( Fn) by thedof something, and so does not change its cohomology class, or its integral o ver a cycle. At first sight, this invariance under deformation suggests t hat all the tr(Fn) are exact forms — they can apparently all be written as tr( Fn) = dω2n−1(A) for some (2 n−1)-formω2n−1(A). To findω2n−1(A) all we have to do is deform the connection to zero by setting At=tAand Ft=dAt+A2 t=tdA+t2A2. (4.92) ThenδAt=Aδt, and d dttr(Fn t) =d/braceleftbig ntr(AFn−1 t)/bracerightbig . (4.93) Integrating up from t= 0, we find tr(Fn) =d/braceleftbigg n/integraldisplay1 0tr(AFn−1 t)dt/bracerightbigg . (4.94) For example tr(F2) =d/braceleftbigg 2/integraldisplay1 0tr(A(tdA+t2A2)dt/bracerightbigg =d/braceleftbigg tr/parenleftbigg AdA+2 3A3/parenrightbigg/bracerightbigg . (4.95) You should recognize here the ω3(A) = tr(AdA+2 3A3) Chern-Simons form of exercise 2.19. The na¨ ıve conclusion — that all the tr( Fn) are exact — is false, however. What the computation actually shows is that when/integraltext tr(Fn)/negationslash= 0 we cannot find a globally defined one-form Arepresenting the connection or gauge field. With no global A, we cannot globally deform Ato zero. Consider, for example, an Abelian U(1) gauge field on the two- sphereS2. When the first Chern-number c1=1 2πi/integraldisplay S2F (4.96) is non-zero, there can be no globally defined one-form Asuch thatF= dA. Glance back, however, at figure 4.10 on page 141. There we see that 150 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY the retractability of the spherical caps D±guarantees that there are one- formsA±defined on D±such thatF=dA±inD±. In the cingular region D+∩D−where they are both defined, A+andA−will be related by a gauge transformation. For a U(1) gauge field, the matrix gappearing in the general gauge transformation rule A→Ag≡g−1Ag+g−1dg, (4.97) of exercise 2.20 becomes the phase eiχ∈U(1). Consequently A+=A−+e−iχdeiχ=A−+idχinD+∩D−. (4.98) The U(1) group element eiχis required to be single valued in D+∩D−, but the angleχmay be multivalued. We now write c1as the sum of integrals over the north and south hemispheres of S2, and use Stokes theorem to reduce this sum to a single integral over the hemispheres’ common bo undary, the equator Γ. c1=1 2πi/integraldisplay northF+1 2πi/integraldisplay southF =1 2πi/integraldisplay northdA++1 2πi/integraldisplay southdA− =1 2πi/integraldisplay ΓA+−1 2πi/integraldisplay ΓA− =1 2π/integraldisplay Γdχ (4.99) We see that c1is the integer counting the winding of χas we circle Γ. An integer cannot be continuously reduced to zero, and if we att empt to deform A→tA→0, we will violate the required single-valuedness of the U(1 ) group elementeiχ. Although the Chern-Simons forms ω2n−1(A) cannot be defined globally, they are still very useful in physics. They occur as Wess-Zumino terms describing the low energy properties of various quantum fiel d theories, the prototype being the Skyrme-Witten model of Hadrons.2 2E. Witten, Nucl. Phys. B223 (1983) 422; ibid.B223 (1983) 433. 4.6. CHARACTERISTIC CLASSES 151 4.6.2 Chern characters and Chern classes Any gauge-invariant polynomial (with exterior multiplica tion of forms un- derstood) in Fprovides a closed, topologically invariant, differential f orm. Certain combinations, however, have additional desirable properties, and so have been given names. The form chn(F) = tr/braceleftbigg1 n!/parenleftbiggi 2πF/parenrightbiggn/bracerightbigg (4.100) is called the n-thChern character . It is convenient to think of this 2 n-form as being the n-th term in a generating-function expansion ch(F)def= tr/braceleftbigg exp/parenleftbiggi 2πF/parenrightbigg/bracerightbigg = ch 0(F) + ch 1(F) + ch 2(F) +···,(4.101) where ch 0(F)≡trIis the dimension of the space on which the ˆλaact. This formal sum of forms of different degree is called the total Chern character . Then! normalization is chosen because it makes the Chern charact er behave nicely when we combine vector bundles. Given two vector bundles over the same manifold, having fibre sUxandVx over the point x, we can make a new bundle with the direct sum Ux⊕Vxas fibre overx. This resulting bundle is called the Whitney sum of the bundles. Similarly we can make a tensor-product bundle whose fibre ove rxisUx⊗Vx. Let us use the notation ch( U) to represent the Chern character of the bundle with fibres Ux, andU⊕Vto denote the Whitney sum. Then we have ch(U⊕V) = ch(U) + ch(V), (4.102) and ch(U⊗V) = ch(U)∧ch(V). (4.103) The second of these formulæ comes about because if ˆλ(1) ais a Lie algebra element acting on V(1)andˆλ(2) athe corresponding element acting on V(2), then they act on the tensor product V(1)⊗V(2)as ˆλ(1⊗2) a=ˆλ(1) a⊗I+I⊗ˆλ(2) a, (4.104) whereIis the identity operator, and for matrices A,B, tr{exp (A⊗I+I⊗B)}= tr{expA⊗expB}= tr{expA}tr{expB}. (4.105) 152 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY In terms of the individual ch n(V) equations (4.102) and (4.103) read chn(U⊕V) = chn(U) + chn(V), (4.106) and chn(U⊗V) =n/summationdisplay m=0chn−m(U)∧chm(V). (4.107) Related to the Chern characters are the Chern classes . These are wedge- product polynomials in the Chern characters, and are defined ,viathe matrix expansion det (I+A) = 1 + trA+1 2/parenleftBig (trA)2−trA2/parenrightBig +..., (4.108) by the generating function for the total Chern class c(F) = det/parenleftbigg I+i 2πF/parenrightbigg = 1 +c1(F) +c2(F) +···. (4.109) Thus c1(F) = ch 1(F), c 2(F) =1 2ch1(F)∧ch1(F)−ch2(F), (4.110) and so on. For matrices AandBwe have det( A⊕B) = det(A) det(B), and this leads to c(U⊕V) =c(U)∧c(V). (4.111) Although the Chern classes are more complicated in appearan ce than the Chern characters, they are introduced because their integr als over cycles are integers , and this property remains true of integer-coefficient sums o f prod- ucts of Chern-classes. The cohomology classes [ cn(F)] are therefore elements of the integer cohomology ring H•(M,Z). This property does not hold for the Chern characters, whose integrals over cycles can be fra ctions. The co- homology classes [ch n(F)] are therefore only elements of H•(M,Q). When we integrate products of Chern classes of total degree 2 mover a closed 2m-dimensional orientable manifold we get integer Chern numbers . These integers can be related to generalized winding number s, and character- ize the extent to which the gauge transformations that relat e the connection fields in different patches serve to twist the vector bundle. Unfortunately it requires a considerable amount of machinery (the Schuber t calculus of complex Grassmannians) to explain these integers. 4.6. CHARACTERISTIC CLASSES 153 Pontryagin and Euler classes When the fibres of a vector bundle are vector spaces over R, the complex skew-hermitian matrices iˆλaare replaced by real skew symmetric matrices. The Lie algebra of the n-by-nmatricesiˆλawas a subalgebra of u(n). The Lie algebra of the n-by-nreal, skew symmetric, matrices is a subalgebra of o(n). Now the trace of an odd power of any skew symmetric matrix is ze ro. As a consequence, Chern characters and Chern classes containin g an odd number ofF’s all vanish. The remaining real 4 n-forms are known as Pontryagin classes . The precise definition is pk(V) = (−1)kc2k(V). (4.112) Pontryagin classes help to classify bundles whose gauge tra nsformations are elements of O( n). If we restrict ourselves to gauge transformations that li e in SO(n), as we would when considering the tangent bundle of an orientable Riemann manifold, then we can make a gauge-invariant polyno mial out of the skew-symmetric matrix-valued Fby forming its Pfaffian . Recall (or see exercise ??.??) that the Pfaffian of a skew symmetric 2 n- by-2nmatrix Awith entries aijis PfA=1 2nn!/epsilon1i1,...i2nai1i2···ai2n−1i2n. (4.113) TheEuler class of the tangent bundle of a 2 n-dimensional orientable manifold is defined viaits skew-symmetric Riemann-curvature form R=1 2Rab,µνdxµdxν(4.114) to be e(R) = Pf/parenleftbigg1 2πR/parenrightbigg . (4.115) In four dimensions, for example, this becomes the 4-form e(R) =1 32π2/epsilon1abcdRabRcd. (4.116) The generalized Gauss-Bonnet theorem asserts, for an oriented, even-dimensional, manifold without boundary, that the Euler character is give n by χ(M) =/integraldisplay Me(R). (4.117) 154 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY We will not prove this theorem, but in section 7.3.6 we will il lustrate the strategy that leads to Chern’s influential proof. Exercise 4.5 : Show that c3(F) =1 6/parenleftBig (ch1(F))3−6ch1(F)ch2(F) + 12ch 3(F)/parenrightBig . 4.7 Hodge Theory and the Morse Index The Laplacian, when acting on a scalar function φinR3is simply div (grad φ), but when acting on a vector vit becomes ∇2v= grad(div v)−curl (curl v). (4.118) Is there a general construction that would have allowed us to write down this second expression? What about the Laplacian on other types o f fields? The Laplacian acting on any vector or tensor field TinRnis given, in general curvilinear co-ordinates, by ∇2T=gµν∇µ∇νTwhere∇µis the flat-space covariant derivative. This is the unique co-ordi nate independent object that reduces in Cartesian co-ordinates to the ordina ry Laplacian acting on the individual components of T. The proof that the rather different- seeming (4.118) holds for vectors is that it too is construct ed out of co- ordinate independent operations and in Cartesian co-ordin ates reduces to the ordinary Laplacian acting on the individual components ofv. It must therefore coincide with the covariant derivative definitio n. Why it should work out this way is not exactly obvious. Now div, grad and cur l can all be expressed in differential form language, and therefore so ca n the scalar and vector Laplacian. Moreover, when we let the Laplacian act on anyp-form the general pattern becomes clear. The differential form defi nition of the Laplacian, and the exploration of its consequences, was the work of William Hodge in the 1930’s. His theory has natural applications to t he topology of manifolds. 4.7.1 The Laplacian on p-forms Suppose that Mis an oriented, compact, D-dimensional manifold without boundary. We can make the space Ωp(M) ofp-form fields on Minto anL2 4.7. HODGE THEORY AND THE MORSE INDEX 155 Hilbert space by introducing the positive-definite inner pr oduct /angbracketlefta,b/angbracketrightp=/angbracketleftb,a/angbracketrightp=/integraldisplay Ma⋆b=1 p!/integraldisplay dDx√gai1i2...ipbi1i2...ip. (4.119) Here the subscript pdenotes the order of the forms in the product, and should not to be confused with the pwe have elsewhere used to label the norm in LpBanach spaces. The presence of the√gand the Hodge ⋆operator tells us that this inner product depends on both the metric on Mand the global orientation. We can use our new product to define a “hermitian adjoint” δ≡d†of the exterior differential operator d. The “...” are because this is not quite an adjoint operator in the normal sense — dtakes us from one vector space to another — but it is constructed in an analogous manner. We d efineδby requiring that /angbracketleftda,b/angbracketrightp+1=/angbracketlefta,δb/angbracketrightp, (4.120) whereais an arbitrary p-form andband arbitrary ( p+ 1)-form. Now recall that⋆takesp-forms to (D−p) forms, and so d⋆bis a (D−p) form. Acting twice on a ( D−p)-form with ⋆gives us back the original form multiplied by (−1)p(D−p). We use this to compute d(a⋆b) =da⋆b + (−1)pa(d⋆b) =da⋆b + (−1)p(−1)p(D−p)a⋆(⋆d⋆b) =da⋆b−(−1)Dp+1a⋆(⋆d⋆b ). (4.121) In obtaining the last line we have observed that p(p−1) is an even integer and so (−1)p(1−p)= 1. Now, using Stokes’ theorem, and the absence of a boundary to discard the integrated-out part, we conclude th at /integraldisplay M(da)⋆b= (−1)Dp+1/integraldisplay Ma⋆(⋆d⋆b ), (4.122) or /angbracketleftda,b/angbracketrightp+1= (−1)Dp+1/angbracketlefta,(⋆d⋆)b/angbracketrightp (4.123) and soδb= (−1)Dp+1(⋆d⋆)b. This was for δacting on a ( p−1) form. Acting on apform we have δ= (−1)Dp+D+1⋆d⋆. (4.124) 156 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY Observe how the sequence of maps in ⋆d⋆works: Ωp(M)⋆−→ΩD−p(M)d−→ΩD−p+1(M)⋆−→Ωp−1(M). (4.125) The net effect is that δtakes ap-form to a ( p−1)-form. Observe also that δ2∝⋆d2⋆= 0. We now define a second-order partial differential operator ∆ pto be the combination ∆p=δd+dδ, (4.126) acting onp-forms This maps a p-form to ap-form. A slightly tedious calcu- lation in cartesian co-ordinates will show that, for flat spa ce, ∆p=−∇2(4.127) on each component of a p-form. This ∆ pis therefore the natural definition for (minus) the Laplacian acting on differential forms. It is usually called the Laplace-Beltrami operator. Using/angbracketlefta,db/angbracketright=/angbracketleftδa,b/angbracketrightwe have /angbracketleft(δd+dδ)a,b/angbracketrightp=/angbracketleftδa,δb/angbracketrightp−1+/angbracketleftda,db/angbracketrightp+1=/angbracketlefta,(δd+dδ)b/angbracketrightp,(4.128) and so we deduce that ∆ pis self-adjoint on Ωp(M). The middle terms in (4.128) are both positive, so we also see that ∆ pis a positive operator — i.e. all its eigenvalues are positive or zero. Suppose that ∆ pa= 0, then (4.128) for a=bbecomes that 0 =/angbracketleftδa,δa/angbracketrightp−1+/angbracketleftda,da/angbracketrightp+1. (4.129) Because both these inner products are positive or zero, the v anishing of their sum requires them to be individually zero. Thus ∆ pa= 0 implies that da=δa= 0. By analogy with harmonic functions, we call a form that is annihilated by ∆ paharmonic form . Recall that a form ais closed ifda= 0. We correspondingly say that aisco-closed ifδa=0. A differential form is therefore harmonic if and only if it is both closed and co-clo sed. When a self-adjoint operator Ais Fredholm ( i.ethe solutions of the equa- tionAx=yare governed by the Fredholm alternative) the vector space o n which it acts is decomposed into a direct sum of the kernel and range of the operator V= Ker(A)⊕Im (A). (4.130) 4.7. HODGE THEORY AND THE MORSE INDEX 157 It may be shown that our Laplace-Beltrami ∆ pis a Fredholm operator, and so for anyp-formωthere is an ηsuch thatωcan be written as ω= (dδ+δd)η+γ =dα+δβ+γ, (4.131) whereα=δη,β=dη, andγis harmonic. This result is known as the Hodge decomposition ofω. It is a form-language generalization of the of the Hodge-Weyl and Helmholtz-Hodge decompositions of chapter ??. It is easy to see that α,βandγare uniquely determined by ω. If they were not then we could find some α,βandγsuch that 0 =dα+δβ+γ (4.132) with non-zero dα,δβandγ. To see that this is not possible, take the dof (4.132) and then the inner product of the result with β. Becaused(dα) = dγ= 0, we end up with 0 =/angbracketleftβ,dδβ/angbracketright =/angbracketleftδβ,δβ/angbracketright. (4.133) Thusδβ= 0. Now apply δto the two remaining terms of (4.132) and take an inner product with α. Becauseδγ= 0, we find/angbracketleftdα,dα/angbracketright= 0, and so dα= 0. What now remains of (4.132) asserts that γ= 0. Suppose that ωis closed. Then our strategy of taking the dof the de- composition ω=dα+δβ+γ, (4.134) followed by an inner product with βleads toδβ= 0. A closed form can thus be decomposed as ω=dα+γ (4.135) withαandγunique. Each cohomology class in Hp(M) therefore contains a unique harmonic representative. Since any harmonic funct ion is closed, and hence a representative of some cohomology class, we conc lude that there is a 1-1 correspondence between p-form solutions of Laplace’s equation and elements of Hp(M). In particular dim(Ker ∆ p) = dim (Hp(M)) =bp. (4.136) 158 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY Herebpis thep-th Betti number. From this we immediately deduce that χ(M) =D/summationdisplay p=0(−1)pdim(Ker ∆ p), (4.137) whereχ(M) is the Euler character of M. There is therefore an intimate relationship between the null-spaces of the second-order p artial differential operators ∆ pand the global topology of the manifold in which they live. This is an example of an index theorem . Just as for the ordinary Laplace operator, ∆ phas a complete set of eigen- functions with associated eigenvalues λ. Because the the manifold is compact and hence has finite volume, the spectrum will be discrete. Re markably, the topological influence we uncovered above is restricted to th e zero-eigenvalue spaces. Suppose that we have a p-form eigenfunction uλfor ∆p: ∆puλ=λuλ. (4.138) Then λduλ=d∆puλ =d(dδ+δd)uλ = (dδ)duλ = (δd+dδ)duλ = ∆p+1duλ. (4.139) Thus, provided it is not identically zero, duλis an (p+1)-form eigenfunction of ∆ (p+1)with eigenvalue λ. Similarly, δuλis a (p−1)-form eigenfunction also with eigenvalue λ. Canduλbe zero? Yes! It will certainly be zero if uλitself is the dof something. What is less obvious is that it will be zero onlyif it is thedof something. To see this suppose that duλ= 0 andλ/negationslash= 0. Then λuλ= (δd+dδ)uλ=d(δuλ). (4.140) Thusduλ= 0 implies that uλ=dη, whereη=δuλ/λ. We see that for λ non-zero, the operators dandδmap theλeigenspaces of ∆ into one another, and the kernel of dacting onp-form eigenfunctions is precisely the image of dacting on (p−1)-form eigenfunctions. In other words, when restricted to positiveλeigenspaces of ∆, the cohomology is trivial. 4.7. HODGE THEORY AND THE MORSE INDEX 159 The set of spaces Vλ ptogether with the maps d:Vλ p→Vλ p+1therefore constitute an exact sequence when λ/negationslash= 0, and so the alternating sum of their dimension must be zero. We have therefore established that /summationdisplay p(−1)pdimVλ p=/braceleftbigg χ(M), λ= 0, 0, λ/negationslash= 0.(4.141) All the topology resides in the null-spaces, therefore. Exercise 4.6 : Show that if ωis closed and co-closed then so is ⋆ω. Deduce that in a for a compact orientable D-manifold we have bp=bD−p. This observation therefore gives another way of understanding P oincar´ e duality. 4.7.2 Morse Theory Suppose, as in the previous section, Mis aD-dimensional compact manifold without boundary and V:M→Ra smooth function. The global topology ofMimposes some constraints on the possible maxima, minima and saddle points ofV. Suppose that P is a stationary point of V. Taking co-ordinates such that P is at xµ= 0, we can expand V(x) =V(0) +1 2Hµνxµxν+.... (4.142) Here, the matrix Hµνis the Hessian Hµν=∂2V ∂xµ∂xν/vextendsingle/vextendsingle/vextendsingle/vextendsingle 0. (4.143) We can change co-ordinates so as reduce the Hessian to a canon ical form with only±1,0 on the diagonal: Hµν= −Im In 0D−m−n . (4.144) If there are no zero’s on the diagonal then the stationary poi nts is said to be non-degenerate . The the number mof downward-bending directions is then called the index ofVat P. If P were a local maximum, then m=D,n= 0. If it were a local minimum then m= 0,n=D. When all its stationary points are non-degenerate, Vis said to be a Morse function . This is the 160 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY generic case. Degenerate stationary points can be regarded as arising from the merging of two or more non-degenerate points. TheMorse index theorem asserts that if Vis a Morse function, and if we defineN0to be the number of stationary points with index 0 ( i.e.local minima), and N1to be the number of stationary points with index 1 etc., then D/summationdisplay m=0(−1)mNm=χ(M). (4.145) Hereχ(M) is the Euler character of M. Thus, a function on the two- dimensional torus, which has χ= 0, can have a local maximum, a local minimum and two saddle points, but cannot have only one local maximum, one local minimum and no saddle points. On a two-sphere ( χ= 2), ifVhas one local maximum and one local minimum it can have no saddle p oints. Closely related to the Morse index theorem is the Poincar´ e-Hopf theorem. It counts the isolated zeros of a tangent-vector field Xon a compact D- manifold and, among other things, explains why we cannot com b a hairy ball. An isolated zero is a pointznat whichXbecomes zero, and that has a neighbourhood in which there is no other zero. If there are on ly finitely many zeros then each of them will be isolated. We can define a vector field index at znby surrounding it with a small ( D−1)-sphere on which Xdoes not vanish. The direction of Xat each point on this sphere then provides a map from the sphere to itself. The index i(zn) is defined to be the winding number (Brouwer degree) of this map. The index can be any integer, but in the sp ecial case thatXis the gradient of a Morse function we have i(zn) = (−1)mnwherem is the Morse index at zn. a) b) c) Figure 4.13: Two-dimensional vector-fields and their streamlines near z eros with indices a) i(za) = +1 , b)i(zb) =−1, c)i(zc) = +1 . 4.7. HODGE THEORY AND THE MORSE INDEX 161 The Poincar´ e-Hopf theorem now states that, for a compact ma nifold with- out boundary, and for a tangent vector field with only finitely many zeros, /summationdisplay zerosni(zn) =χ(M). (4.146) A tangent-vector field must therefore always have at least on e zero unless χ(M) = 0. Since the two-sphere has χ= 2, it cannot be combed. Figure 4.14: Gradient vector field and streamilines in a two-simplex. If one is prepared to believe that/summationtext zerosi(zn) is the same integer for all tangent vector fields XonM, it is simple to show that this integer must be equal to the Euler character of M. Consider, for ease of visualization, a two-manifold. Triangulate Mand takeXto be the gradient field of a function with local minima at each vertices, saddle points o n the edges, and local maxima at the centre of each face (see figure 4.14). It mu st be clear that this particular field Xhas /summationdisplay zerosni(zn) =V−E+F=χ(M). (4.147) In the case of a two-dimensional oriented surface equipped w ith a smooth metric, it is also simple to demonstrate the invariance of th e index sum. Consider two vector fields XandY. Triangulate Mso that all zeros of both fields lie in the interior of the faces of the simplices. The me tric allows us to compute the angle θbetweenXandYwherever they are both non-zero, and in particular on the edges of the simplices. For each two- simplexσwe compute the total change ∆ θin the angle as we circumnavigate its boundary. This change is an integral multiple of 2 π, with the integer counting the difference /summationdisplay zeros ofX∈σi(zn)−/summationdisplay zeros ofY∈σi(zn) (4.148) 162 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY of the indices of the zeros within σ. On summing over all triangles σ, each edge is traversed twice, once in each direction, so/summationtext σ∆θvanishes . The total index ofXis therefore the same as that of Y. This pairwise cancellation argument can be extended to non- orientable surfaces, such as the projective plane, In this case the edge s constituting the homological “boundary” of the closed surface are traversed twice in the same direction, but the angle θat a point on one edge is paired with −θat the corresponding point of the other edge. Supersymmetric Quantum Mechanics Edward Witten gave a beautiful proof of the Morse index theor em for an orientable manifold by re-interpreting the Laplace-Beltr ami operator as the Hamiltonian of supersymmetric quantum mechanics onM. Witten’s idea had a profound impact, and led to quantum physics serving as a ric h source of inspiration and insight for mathematicians. We have seen mo st of the ingre- dients of this re-interpretation in previous chapters. Ind eed you should have experienced a sense of d´ ej` a vu when you saw dandδmapping eigenfunctions of one differential operator into eigenfunctions of a relate d operator. We begin with an novel way to think of the calculus of different ial forms. We introduce a set of fermion annihilation and creation oper atorsψµand ψ†µwhich anti-commute, ψµψν=−ψνψµ, and obey {ψ†µ,ψν}≡ψ†µψν+ψνψ†µ=gµν. (4.149) Hereµruns from 1 to D. As is usual when we are given such operators, we also introduce a vacuum state|0/angbracketrightwhich is killed by all the annihilation operators:ψµ|0/angbracketright= 0. The states (ψ†1)p1(ψ†2)p2...(ψ†n)pn|0/angbracketright, (4.150) with each of the pitaking the value one or zero, then constitute a basis for 2D-dimensional space. We call p=/summationtext ipithefermion number of the state. We now assume that /angbracketleft0|0/angbracketright= 1 and use the anti-commutation relations to show that /angbracketleft0|ψµp...ψµ2ψµ1...ψ†ν1ψ†ν2...ψ†νq|0/angbracketright is zero unless p=q, in which case it is equal to gµ1ν1gµ2ν2...gµpνp±(permutations) . 4.7. HODGE THEORY AND THE MORSE INDEX 163 We now make the correspondence 1 p!fµ1µ2...µp(x)ψ†µ1ψ†µ2...ψ†µp|0/angbracketright↔1 p!fµ1µ2...µp(x)dxµ1dxµ2...dxµp, (4.151) to identify p-fermion states with p-forms. We think of fµ1µ2...µp(x) as being the wavefunction of a particle moving on M, with the subscripts informing us there are fermions occupying the states µi. It is then natural to take the inner product of |a/angbracketright=1 p!aµ1µ2...µp(x)ψ†µ1ψ†µ2...ψ†µp|0/angbracketright (4.152) and |b/angbracketright=1 q!bµ1µ2...µq(x)ψ†µ1ψ†µ2...ψ†µq|0/angbracketright (4.153) to be /angbracketlefta,b/angbracketright=/integraldisplay MdDx√g1 p!q!a∗ µ1µ2...µpbν1ν2...νq/angbracketleft0|ψµp...ψµ1ψ†ν1...ψ†νq|0/angbracketright =δpq/integraldisplay MdDx√g1 p!a∗ µ1µ2...µpbµ1µ2...µp. (4.154) This coincides the Hodge inner product of the corresponding forms. If we lower the index by setting ψµto begµνψµthen the action of Xµψµ on ap-fermion state coincides with the action of the interior mul tiplication iXon the corresponding p-form. All the other operations of the exterior calculus can also be expressed in terms of the ψ’s. In particular, in Cartesian co-ordinates where gµν=δµν, we can identify dwithψ†µ∂µ. To find the operator that corresponds to the Hodge δ, we compute δ=d†= (ψ†µ∂µ)†=∂† µψµ=−∂µψµ=−ψµ∂µ. (4.155) The hermitian adjoint of ∂µis here being taken with respect to the standard L2(RD) inner product. This computation becomes more complicated when whengµνbecomes position dependent. The adjoint ∂† µthen involves the derivative of√g, andψand∂µno longer commute. For this reason, and because such complications are inessential for what follow s, we will delay discussing this general case until the end of this section. Having found a simple formula for δ, it is now automatic to compute dδ+δd=−{ψ†µ,ψν}∂µ∂ν=−δµν∂µ∂ν=−∇2. (4.156) 164 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY This much easier than deriving the same result by using δ= (−1)Dp+D+1⋆d⋆. Witten’s fermionic formalism simplifies a number of compuat ions involv- ingδ, but his real innovation was to consider a deformation of the exterior calculus by introducing the operators dt=e−tV(x)detV(x), δt=etV(x)δe−tV(x), (4.157) and ∆t=dtδt+δtdt. (4.158) HereV(x) is the Morse function whose stationary points we are seekin g to count. The deformed derivative continues to obey d2 t= 0, anddω= 0 if and only ifdte−tVω= 0. Similarly, if ω=dηthene−tVω=dte−tVη. The cohomol- ogy ofdanddtare therefore transformed into each other by multiplicatio n bye−tV. Since the exponential function is never zero, this corresp ondence is invertible and the mapping is an isomorphism. In particul ar, the Betti numbersbp, the dimensions of Ker ( dt)p/Im (dt)p−1, aretindependent. Fur- ther, thet-deformed Laplace-Beltrami operator remains Fredholm wit h only positive or zero eigenvalues. We can make a Hodge decomposit ion ω=dtα+δtβ+γ, (4.159) where ∆ tγ= 0, and concude that dim (Ker (∆ t)p) =bp (4.160) as before. The non-zero eigenvalue spaces will also continu e to form exact sequences. Nothing seems to have changed! Why do we introduc edtthen? The motivation is that when tbecomes large we can use our knowledge of quantum mechanics to compute the Morse index. To do this, we expand out dt=ψ†µ(∂µ+t∂µV) δt=−ψµ(∂µ−t∂µV) (4.161) and find dtδt+δtdt=−∇2+t2|∇V|2+t[ψ†µ,ψν]∂2 µνV. (4.162) This can be thought of as a Schr¨ odinger Hamiltonian on Mcontaining a potential and a fermionic term. When tis large and positive the potential 4.7. HODGE THEORY AND THE MORSE INDEX 165 t2|∇V|2will be large everywhere except near those points where ∇V= 0. The wavefunctions of all low-energy states, and in particul ar all zero-energy states, will therefore be concentrated at precisely the sta tionary points we are investigating. Let us focus on a particular stationary poin t, which we will take as the origin of our co-ordinate system, and identify an y zero-energy state localized there. We first rotate the coordinate system about the origin so that the Hessian matrix ∂2 µνV|0becomes diagonal with eigenvalues λn. The Schr¨ odinger problem can then be approximated by a sum of harmonic oscillator hamiltonians ∆p,t≈D/summationdisplay i=1/braceleftbigg −∂2 ∂x2 i+t2λ2 ix2 i+tλi[ψ†i,ψi]/bracerightbigg . (4.163) The commutator [ ψ†i,ψi] takes the value +1 if the i’th fermion state is oc- cupied, and−1 if it is not. The spectrum of the approximate Hamiltonian is therefore tD/summationdisplay i=1{|λi|(1 + 2ni)±λi}. (4.164) Here thenilabel the harmonic oscillator states. The lowest energy sta tes will have all the ni= 0. To get a state with zero energy we must arrange for the±sign to be negative (no fermion in state i) whenever λiis positive, and to be positive (fermion state ioccupied) whenever λiis negative. The fermion number “ p” of the zero-energy state is therefore equal to the the number of negative λi—i.e.to the index of the critical point! We can, in this manner, find one zero-energy state for each critical p oint. All other states have energies proportional t, and therefore large. Since the number of zero energy states having fermion number pis the Betti number bp, the harmonic oscillator approximation suggests that bp=Np. If we could trust our computation of the energy spectrum, we w ould have established the Morse theorem D/summationdisplay p=0(−1)pNp=D/summationdisplay p=0(−1)pbp=χ(M), (4.165) by having the two sums agree term by term. Our computation is o nly ap- proximate, however. While there can be no more zero-energy s tates than those we have found, some states that appear to be zero modes m ay instead 166 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY have small positive energy. This might arise from tunnellin g between the different potential minima, or from the higher-order correc tions to the har- monic oscillator potentials, both effects we have neglected . We can therefore only be confident that Np≥bp. (4.166) The remarkable thing is that, for the Morse index, this does not matter ! If one of our putative zero modes gains a small positive energy, it is now in the non-zero eigenvalue sector of the spectrum. The exact-s equence property therefore tells us that one of the other putative zero modes m ust also be a not-quite-zero mode state with exactly the same energy. Thi s second state will have a fermion number that differs from the first by plus or minus one. Our error in counting the zero energy states therefore cance ls out when we take the alternating sum. Our unreliable estimate bp≈Nphas thus provided us with an exact computation of the Morse index. We have described Witten’s argument as if the manifold Mwere flat. When the manifold Mis not flat, however, the curvature will not affect our computations. Once the parameter tis large the low-energy eigenfunc- tions will be so tightly localized about the critical points that they will be hard-pressed to detect the curvature. Even if the curvature can effect an infintesimal energy shift, the exact-sequence argument aga in shows that this does not affect the alternating sum. The Weitzenb¨ ock Formula Although we we were able to evade them when proving the Morse i ndex theorem, it is interesting to uncover the workings of the nit ty-gritty Rie- mann tensor index machinary that lie concealed behind the po lished facade of Hodge’s d,δcalculus. Let us assume that our manifold Mis equipped with a torsion-free con- nection Γµνλ= Γµλν, and use this connection to define the action of an operator ˆ∇µby specifying its commutators with c-number functions f, and with theψµandψ†µ’s: [ˆ∇µ,f] =∂µf, [ˆ∇µ,ψ†ν] =−Γν µλψ†λ, [ˆ∇µ,ψν] =−Γν µλψλ. (4.167) 4.7. HODGE THEORY AND THE MORSE INDEX 167 We also set ˆ∇µ|0/angbracketright= 0. These rules allow us to compute the action of ˆ∇µon fµ1µ2...µp(x)ψ†µ1...ψ†µp|0/angbracketright. For example ˆ∇µ/parenleftbig fνψ†ν|0/angbracketright/parenrightbig =/parenleftBig [ˆ∇µ,fνψ†ν] +fνψ†νˆ∇µ/parenrightBig |0/angbracketright =/parenleftBig [ˆ∇µ,fν]ψ†ν+fα[ˆ∇µ,ψ†α]/parenrightBig |0/angbracketright = (∂µfν−fαΓα µν)ψ†ν|0/angbracketright = (∇µfν)ψ†ν|0/angbracketright, (4.168) where ∇µfv=∂µfν−Γα µνfα, (4.169) is the usual covariant derivative acting on the componenent s of a covariant vector. The metric gµνcounts as a c-number function, and so [ ˆ∇α,gµµ] is not zero, but is instead ∂αgµν. This might be disturbing—being able pass the metric through a covariant derivative is a basic compatibil ty condition in Riemann geometry—but all is not lost. ˆ∇µ(with a caret) is not quite the same beast as∇µ. We proceed as follows: ∂αgµν= [ˆ∇α,gµµ] = [ˆ∇α,{ψ†µ,ψν}] = [ˆ∇α,ψ†µψν] + [ˆ∇α,ψνψ†µ,] =−{ψ†µ,ψλ}Γν αλ−{ψ†ν,ψλ}Γµ αλ =−gµλΓν αλ−gνλΓµ αλ. (4.170) We conclude that ∂αgµν+gµλΓν αλ+gλνΓµ αλ≡∇αgµν= 0. (4.171) Metric compatibility is therefore satisfied, and the connec tion is therefore the standard Riemannian Γα µν=1 2gαλ(∂µgλν+∂νgµλ−∂λgµν). (4.172) Knowing this, we can compute the adjoint of ˆ∇µ: /parenleftBig ˆ∇µ/parenrightBig† =−1√gˆ∇µ√g =−/parenleftBig ˆ∇µ+∂µln√g/parenrightBig =−(ˆ∇µ+ Γν νµ). (4.173) 168 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY That Γννµis the logarithmic derivative of√gis a standard identity for the Riemann connection (see exercise 2.14). The resultant form ula for ( ˆ∇µ)† can be used to verify that the second and third equations in (4 .167) are compatible with each other. We can also compute [[ ˆ∇µ,ˆ∇ν],ψα] and from it deduce that [ˆ∇µ,ˆ∇ν] =Rσλµνψ†σψλ, (4.174) where Rα βµν=∂µΓα βν−∂νΓα βµ+ Γα λµΓλ βν−Γα λνΓλ βµ (4.175) is the Riemann curvature tensor. We now define dto be d=ψ†µˆ∇µ. (4.176) Its action coincides with the usual dbecause the symmetry of the Γα µν’s ensures that their contributions cancel. From this we find th atδis δ≡/parenleftBig ψ†µˆ∇µ/parenrightBig† =ˆ∇† µψµ =−(ˆ∇µ+ Γν µν)ψµ =−ψµ(ˆ∇µ+ Γν µν) + Γµ µνψν =−ψµˆ∇µ. (4.177) The Laplace-Beltrami operator can now be worked out as dδ+δd=−/parenleftBig ψ†µˆ∇µψνˆ∇ν+ψνˆ∇νψ†µˆ∇µ/parenrightBig =−/parenleftBig {ψ†µ,ψν}(ˆ∇µˆ∇ν−Γσ µνˆ∇σ) +ψνψ†µ[ˆ∇ν,ˆ∇µ]/parenrightBig =−/parenleftBig gµν(ˆ∇µˆ∇ν−Γα µνˆ∇σ) +ψνψ†µψ†σψλRσλνµ/parenrightBig (4.178) By making use of the symmetries Rσλνµ=RνµσλandRσλνµ=−Rσλµνwe can tidy up the curvature term to get dδ+δd=−gµν(ˆ∇µˆ∇ν−Γσ µνˆ∇σ)−ψ†αψβψ†µψνRαβµν. (4.179) This result is called the Weitzenb¨ ock formula . An equivalent formula can be derived directly from (4.124), but only with a great deal mor e effort. The part 4.7. HODGE THEORY AND THE MORSE INDEX 169 without the curvature tensor is called the Bochner Laplacian . It is normally written asB=−gµν∇µ∇νwith∇µbeing understood to be acting on the indexν, and therefore tacitly containing the extra Γσ µνthat must be made explicit when we define the action of ˆ∇µviacommutators. The Bochner Laplacian can also be written as B=ˆ∇† µgµνˆ∇ν (4.180) which shows that it is a positive operator. 170 CHAPTER 4. AN INTRODUCTION TO TOPOLOGY Chapter 5 Groups and Group Representations Groups appear in physics as symmetries of the system we are st udying. Often the symmetry operation involves a linear transformation, a nd this naturally leads to the idea of finding sets of matrices having the same mu ltiplication table as the group. These sets are called representations of the group. Given a group, we endeavour to find and classify all possible repres entations. 5.1 Basic Ideas We begin with a rapid review of basic group theory. 5.1.1 Group Axioms AgroupGis a set with a binary operation that assigns to each ordered p air (g1,g2) of elements a third element, g3, usually written with multiplicative notation as g3=g1g2. The binary operation, or product , obeys the following rules: i) Associativity: g1(g2g3) = (g1g2)g3. ii) Existence of an identity: There is an element1e∈Gsuch thateg=g for allg∈G. 1The symbol “ e” is often used for the identity element, from the German Einheit , meaning “unity.” 171 172 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS iii) Existence of an inverse: For each g∈Gthere is an element g−1such thatg−1g=e. From these axioms there follow some conclusions that are so b asic that they are often included in the axioms themselves, but since t hey are not independent, we state them as corollaries. Corollary i) :gg−1=e. Proof : Start from g−1g=e, and multiply on the right by g−1to get g−1gg−1=eg−1=g−1, where we have used the left identity property of eat the last step. Now multiply on the left by ( g−1)−1, and use associativity to getgg−1=e. Corollary ii) :ge=g. Proof : Writege=g(g−1g) = (gg−1)g=eg=g. Corollary iii) : The identity eis unique. Proof : Suppose there is another element e1such thate1g=eg=g. Multiply on the right by g−1to gete1e=e2=e, bute1e=e1, soe1=e. Corollary iv) : The inverse of a given element gis unique. Proof : Letg1g=g2g=e. Use the result of corollary (i), that any left inverse is also a right inverse, to multiply on the right by g−1 1, and so find thatg1=g2. Two elements g1andg2are said to commute ifg1g2=g2g1. If the group has the property that g1g2=g2g1for allg1,g2∈G, it is said to be Abelian , otherwise it is non-Abelian . If the setGcontains only finitely many elements, the group Gis said to befinite. The number of elements in the group, |G|, is called the order of the group. Examples of Groups: 1) The integers Zunder addition. The binary operation is ( n,m)/mapsto→n+m, and “0” plays the role of the identity element. This is not a fin ite group. 2) The integers modulo nunder addition. ( m,m/prime)/mapsto→m+m/prime,modn. This group is denoted by Zn. 3) The non-zero integers modulo p(a prime) under multiplication (m,m/prime)/mapsto→ mm/prime,modp. Here “1” is the identity element. If the modulus is not a prime number, we do not get a group (why not?). This group is sometimes denoted by ( Zp)×. 5.1. BASIC IDEAS 173 4) The set of numbers {2,4,6,8}under multication modulo 10. Here, the number “6” plays the role of the identity! 5) The set of functions f1(z) =z, f 2(z) =1 1−z, f 3(z) =z−1 z f4(z) =1 z, f 5(z) = 1−z, f 6(z) =z z−1 with (fi,fj)/mapsto→fi◦fj. Here the “◦” is a standard notation for compo- sition of functions: ( fi◦fj)(z) =fi(fj(z)). 6) The set of rotations in three dimensions, equivalently th e set of 3-by-3 real matrices O, obeyingOTO=I, and detO= 1. This is the group SO(3). SO( n) is defined analogously as the group of rotations in n dimensions. If we relax the condition on the determinant we g et the orthogonal group O(n). Both SO( n) and O(n) are examples of Lie groups . A Lie group a group that is also a manifold M, and whose multiplication law is a smooth function M×M→M. 7) Groups are often specified by giving a list of generators andrelations . For example the cyclic group of ordern, denoted by Cn, is specified by giving the generator aand relation an=e. Similarly, the dihedral group Dnhas two generators a,band relations an=e,b2=e, (ab)2=e. This group has order 2 n. 5.1.2 Elementary Properties Here are the basic properties of groups that we need: i)Subgroups : If a subset of elements of a group forms a group, it is called a subgroup. For example, Z12has a subgroup of consisting of {0,3,6,9}. All groups have at least two subgroups: the trivial sub- groupsGitself, and{e}. Any other subgroups are called proper sub- groups. ii)Cosets : Given a subgroup H⊆G, having elements {h1,h2,...}, and an element g∈G, we form the (left) cosetgH={gh1,gh2,...}. If two cosetsg1Handg2Hintersect, they coincide. (Proof: if g1h1=g2h2, theng2=g1(h1h−1 2) and sog1H=g2H.) IfHis a finite group, each coset has the same number of distinct elements as H. (Proof: if gh1=gh2then left multiplication by g−1shows that h1=h2.) If the 174 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS order ofGis also finite, the group Gis decomposed into an integer number of cosets, G=g1H+g2H+···, (5.1) where “+”denotes the union of disjoint sets. From this we see that the order ofHmust divide the order of G. This result is called Lagrange’s theorem . The set whose elements are the cosets is denoted by G/H. iii)Normal subgroups and quotient groups : A subgroup HofGis said to be normal , orinvariant , ifg−1Hg=Hfor allg∈G. Given a normal subgroup H, we can define a multiplication rule on the coset space cosets G/H≡{g1H,g2H,...}by taking a representative element from each of giH, andgjH, taking the product of these elements, and defining (giH)(gjH) to be the coset in which this product lies. This coset is independent of the representative elements chosen (this would not be so if the subgroup was not normal). The resulting group is called the quotient group G/H. (Note that the symbol “ G/H” is used to denote both the set of cosets, and, when it exists, the grou p whose elements are these cosets.) iv)Simple groups : A groupGwith no normal subgroups is said to be sim- ple. The finite simple groups have been classified. They fall into various infinite families (Cyclic groups, Alternating groups, 16 fa milies of Lie type) together with 26 sporadic groups , the largest of which, the Mon- ster, has order 808,017,424,794,512,875,886,459,904,961,71 0,757,005, 754, 368,000,000,000. The mysterious “Monstrous moonshine” li nks its rep- resentation theory to the elliptic modular function J(τ) and to string theory. iv)Conjugacy and Conjugacy Classes : Two group elements g1,g2are said to beconjugate inGif there is an element g∈Gsuch thatg2=g−1g1g. Ifg1is conjugate to g2, we writeg1∼g2. Conjugacy is an equivalence relation ,2and, for finite groups, the resulting conjugacy classes have order that divide the order of G. To see this, consider the conjugacy class containing an element g. Observe that the set Hof elements h∈Gsuch thath−1gh=gforms a subgroup. The set of elements 2An equivalence relation, ∼, is a binary relation that is i)Reflexive :A∼A. ii)Symmetric :A∼B⇐⇒B∼A. iii)Transitive :A∼B, B∼C=⇒A∼C Such a relation breaks a set up into disjoint equivalence classes. 5.1. BASIC IDEAS 175 conjugate to gcan be identified with the coset space G/H. The order ofGdivided by the order of the conjugacy class is therefore |H|. Example : In the rotation group SO(3), the conjugacy classes are the s ets of rotations through the same angle, but about different axes. Example : In the group U( n), ofn-by-nunitary matrices, the conjugacy classes are the set of matrices possessing the same eigenval ues. Example: Permutations. The permutation group on nobjects,Sn, has order n!. Suppose we consider permutations π1,π2inS8such thatπ1that maps π1: 1 2 3 4 5 6 7 8 ↓ ↓ ↓ ↓ ↓ ↓ ↓ ↓ 2 3 1 5 4 7 6 8 , andπ2maps π2: 1 2 3 4 5 6 7 8 ↓ ↓ ↓ ↓ ↓ ↓ ↓ ↓ 2 3 4 5 6 7 8 1 . The product π2◦π1then takes π2◦π1: 1 2 3 4 5 6 7 8 ↓ ↓ ↓ ↓ ↓ ↓ ↓ ↓ 3 4 2 6 5 8 7 1 . We can write these partitions out more compactly by using Pao lo Ruffini’s cycle notation: π1= (123)(45)(67)(8) , π 2= (12345678) , π 2◦π1= (132468)(5)(7) . In this notation, each number is mapped to the one immediatel y to its right, with the last number in each bracket, or cycle, wrapping round to map to the first. Thus π1(1) = 2,π1(2) = 3,π1(3) = 1. The “8”, being both first and last in its cycle, maps to itself: π1(8) = 8. Any permutation with this cycle pattern, (∗∗∗)(∗∗)(∗∗)(∗), is in the same conjugacy class as π1. We say thatπ1possesses one 1-cycle, two 2-cycles, and one 3-cycle. The cl ass (r1,r2,...rn) havingr11-cycles,r22-cycles etc., wherer1+2r2+···+nrn=n, contains N(r1,r2,...)=n! 1r1(r1!) 2r2(r2!)···nrn(rn!) elements. The signof the permutation, sgnπ=/epsilon1π(1)π(2)π(3)...π(n) 176 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS is equal to sgnπ= (+1)r1(−1)r2(+1)r3(−1)r4···. We have, for any two permutations π1,π2 sgn (π1)sgn (π2) = sgn (π1◦π2), so the even(sgnπ= +1) permutations form an invariant subgroup called theAlternating group ,An. The group Anis simple for n≥5, and Ruffini (1801) showed that this simplicity prevents the solution of the general quin- tic by radicals. His work was ignored, however, and later ind ependently rediscovered by Abel (1824) and Galois (1829). If we write out the group elements in some order {e,g1,g2,...}, and then multiply on the left g{e,g1,g2,...}={g,gg 1,gg2,...} then the ordered list {g,gg 1,gg2,...}is a permutation of the original list. Any group is therefore a subgroup of S|G|. This is called Cayley’s Theorem . Exercise 5.1 : LetH1,H2be two subgroups of a group G. Show that H1∩H2 is also a subgroup. Exercise 5.2 : LetGbe any group. a) The subset Z(G) ofGconsisting of those g∈Gthat commute with all other elements of the group is called the center of the group. Show that Z(G) is a subgroup of G. b) Ifgis an element of G, the setCG(g) of elements of Gthat commute withgis called the centeralizer ofginG. Show that it is a subgroup of G. c) IfHis a subgroup, the set of elements of Gthat commute with all elements of His the centralizer CG(H) ofHinG. Show that it is a subgroup of G. d) IfHis a subgroup, the set NG(H)⊂Gconsisting of those gsuch that g−1Hg=His called the normalizer ofHinG. Show that NG(H) is a subgroup of G, and thatHis a normal subgroup of NG(H). Exercise 5.3 : Show that the set of powers anof an element a∈Gform a subgroup. Let pbe prime. Recall that the set {1,2,...p−1}forms the group (Zp)×under multiplication modulo p. By appealing to Lagrange’s theorem, prove Fermat’s little theorem that for any prime pand integer a, we have ap−1= 1,modp. 5.1. BASIC IDEAS 177 Exercise 5.4 : Use Fermat’s theorem from the previous excercise to establ ish the mathematical identity underlying the RSA algorithm for public-key cryp- tography: Let p,qbe prime and N=pq. First use Euclid’s algorithm for the HCF of two numbers to show that if the integer eis co-prime to3(p−1)(q−1), then there is an integer dsuch that de= 1,mod(p−1)(q−1). Then show that if, C=Me,modN, (encryption) then M=Cd,modN. (decryption) . The numbers eandNcan be made known to the public, but it is hard to find the secret decoding key, d, unless the factors pandqofNare known. Exercise 5.5 : Consider the group Gwith multiplication table shown in table 5.1. GI A B C D E II A B C D E AA B I E C D BB I A D E C CC D E I A B DD E C B I A EE C D A B I Table 5.1: Multiplication table of G. To findABlook in row AcolumnB. This group has proper a subgroup H={I,A,B}, and corresponding (left) cosets areIH={I,A,B}andCH={C,D,E}. (i) Construct the conjugacy classes of this group. (ii) Show that{I,A,B}and{C,D,E}are indeed the left cosets of H. (iii) Determine whether His a normal subgroup. (iv) If so, construct the group multiplication table for the corresponding quo- tient group. 3Has no factors in common with. 178 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS Exercise 5.6 : LetHandK, be groups. Make the cartesian product G=H×K into a group by introducing a multiplication rule for elemen ts of the Cartesian product by setting: (h1,k1)∗(h2,k2) = (h1h2,k1k2). Show thatG, equipped with∗as its product, satsifies the group axioms. The resultant group is called the direct product ofHandK. Exercise 5.7 : IfFandGare groups, a map ϕ:F→Gthat preserves the group structure, i.e.ifϕ(g1)ϕ(g2) =ϕ(g1g2), is called a group homomorphism. If ϕis such a homomorphism show that ϕ(eF) =eG, whereeF, andeGare the identity element in F,Grespectively. Exercise 5.8 :. Ifϕ:F→Gis a group homomorphism, and if we define Ker( ϕ) as the set of elements f∈Fthat map to eG, show that Ker( ϕ) is a normal subgroup of F. 5.1.3 Group Actions on Sets Groups usually appear in physics as symmetries: they act on a physical object to change it in some way, perhaps while leaving some ot her property invariant. SupposeXis a set. We call its elements “points.” A group action onX is a mapg∈G:X→Xthat takes a point x∈Xto a new point that we denote bygx∈X, and such that g2(g1x) = (g1g2)x, andex=x. There is some standard vocabulary for group actions: i) Given a a point x∈Xwe define the orbit ofxto be the set Gx≡ {gx:g∈G}⊆X. ii) The action of the group is transitive if any orbit is the whole of X. iii) The action is effective , orfaithful , if the map g:X→Xbeing the identity map implies that g=e. Another way of saying this is that the action is effective if the map G→Map (X→X) is one-to-one. If the action of Gisnotfaithful, the set of g∈Gthat act as the identity map forms an invariant subgroup HofG, and the quotient group G/H has a faithful action. iv) The action is freeif the existence of an xsuch thatgx=ximplies that g=e. In this case, we also say that gacts without fixed points. 5.2. REPRESENTATIONS 179 If the group acts freely and transitively, then having chose n a fiducial pointx0, we can uniquely label every point in Xby the group element g such thatx=gx0. (Ifg1andg2both takex0→x, theng−1 1g2x0=x0. By the free action property we deduce that g−1 1g2=e, andg1=g2.). In this case we might, for some purposes, identify XwithG. Suppose the group acts transitively, but not freely. Let Hbe the set of elements that leaves x0fixed. This is clearly a subgroup of G, and if g1x0=g2x0we haveg−1 1g2∈H, org1H=g2H. The space Xcan therefore be identified with the space of cosets G/H. Such sets are called quotient spaces orHomogeneous spaces. Many spaces of significance in physics can be though of as cosets in this way. Example : The rotation group SO(3) acts transitively on the two-sphe reS2. The SO(2) subgroup of rotations about the zaxis, leaves the north pole of the sphere fixed. We can therefore identify S2/similarequalSO(3)/SO(2). Many phase transitions are a result of spontaneous symmetry breaking . For example the water →ice transition results in the continuous translation invariance of the liquid water being broken down to the discr ete translation invariance of the crystal lattice of the solid ice. When a sys tem with symme- try groupGspontaneously breaks the symmetry to a subgroup H, the set of inequivalent ground states can be identified with the homo geneous space G/H. 5.2 Representations Ann-dimensional representation of a group is formally defined to be a homo- morphism from Gto a subgroup of GL( n,C), the group of invertible n-by-n matrices with complex entries. In effect, it is a set of n-by-nmatrices that obeys the group multiplication rules D(g1)D(g2) =D(g1g2), D(g−1) = [D(g)]−1. (5.2) Given such a representation, we can form another one D/prime(g) by conjuga- tion with any fixed invertible matrix C D/prime(g) =C−1D(g)C. (5.3) IfD/prime(g) is obtained from D(g) in this way, we say that they are equivalent representations and write D∼D/prime. We can think of DandD/primeas being 180 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS matrices representing the same linear map, but in different b ases. Our task in the rest of this chapter is to find and classify all represen tations of a finite groupGup to equivalence. Real and pseudo-real representations We can form a new representation from D(g) by setting D/prime(g) =D∗(g), whereD∗(g) denotes the matrix whose entries are the complex conjugate s of those in D(g). Suppose D∗∼D. It may then be possible to find a basis in which the matrices have only real entries. In this ca se we say the representation is real. It may be, however, be that D∗∼Dbut we cannot find a basis in which the matrices become real. In this case we s ay thatDis pseudo-real . Example: Consider the defining representation of SU(2) (the group of 2 -by-2 unitary matrices with unit determinant.) Such matrices are necessarily of the form U=/parenleftbigg a−b∗ b a∗/parenrightbigg , (5.4) whereaandbare complex numbers with |a|2+|b|2= 1. They are there- fore specified by three real parameters, and so the group manifold is three dimensional. Now /parenleftbigg a−b∗ b a∗/parenrightbigg∗ =/parenleftbigg a∗−b b∗a/parenrightbigg , =/parenleftbigg 0 1 −1 0/parenrightbigg/parenleftbigg a−b∗ b a∗/parenrightbigg/parenleftbigg 0−1 1 0/parenrightbigg , =/parenleftbigg 0−1 1 0/parenrightbigg−1/parenleftbigg a−b∗ b a∗/parenrightbigg/parenleftbigg 0−1 1 0/parenrightbigg , (5.5) and soU∼U∗. It is not possible to find a basis in which all SU(2) matrices are simultaneously real, however. If such a basis existed we could specify the matrices by only two real parameters—but we have seen that we need three real numbers to describe all possible SU(2) matrices. 5.2. REPRESENTATIONS 181 Direct Sum and Direct Product We can obtain new representations from old by combining them . Given two representations D(1)(g),D(2)(g), we can form their direct sum D(1)⊕D(2)as the block-diagonal matrix /parenleftbigg D(1)(g) 0 0D(2)(g)/parenrightbigg . (5.6) We are particularly interested in taking a representation a nd breaking it up as a direct sum of irreducible representations. Given two representations D(1)(g),D(2)(g), we can combine them in a different way by taking their direct product D(1)⊗D(2), the natural action of the group on the tensor product of the representation spac es. In other words, ifD(1)(g)e(1) j=e(1) iD(1) ij(g) andD(2)(g)e(2) j=e(2) iD(2) ij(g) we define [D(1)⊗D(2)](g)(e(1) i⊗e(2) j) = (e(1) k⊗e(2) l)D(1) ki(g)D(2) lj(g).(5.7) We think of D(1) ki(g)D(2) lj(g) being the entries in the direct-product matrix matrix [D(1)(g)⊗D(2)(g)]kl,ij, whose rows and columns are indexed by pairs of numbers. The dimension of the product representation is therefore the product of the d imensions of its factors. Exercise 5.9 : Show that if D(g) is a representation, then so is D/prime(g) = [D(g−1)]T, where the superscript Tdenotes the transposed matrix. Exercise 5.10 : Show that a map that assigns every element of a group Gto the 1-by-1 identity matrix is a representation. It is, not un reasonably, called thetrivial representation. Exercise 5.11 : A representation D:G→GL(n,C) that assigns an element g∈Gto then-by-nidentity matrix Inif and only if g=eis said to be faithful . LetDbe a non-trivial, but non-faithful, representation of Gbyn- by-nmatrices. Let H⊂Gconsist of those elements hsuch thatD(h) =In. Show that His a normal subgroup of G, and that Dprojects to a faithful representation of the quotient group G/H. 182 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS Exercise 5.12 : LetAandBbe linear maps from U→UandCandDbe linear maps from V→V. Then the direct products A⊗CandB⊗Dare linear maps from U⊗V→U⊗V. Show that (A⊗C)(B⊗D) = (AB)⊗(CD). Show also that (A⊕C)(B⊕D) = (AB)⊕(CD). Exercise 5.13 : LetAandBbem-by-mandn-by-nmatrices respectively, and letIndenote the n-by-nunit matrix. Show that: i) tr(A⊕B) = tr(A) + tr(B). ii) tr(A⊗B) = tr(A)tr(B). iii) exp(A⊕B) = exp(A)⊕exp(B). iv) exp(A⊗In+Im⊗B) = exp(A)⊗exp(B). v) det(A⊕B) = det(A)det(B). vi) det(A⊗B) = (det(A))n(det(B))m. 5.2.1 Reducibility and Irreducibility The “atoms” of representation theory are those representat ions that cannot, by a clever choice of basis, be decomposed into, or reduced to, a direct sum of smaller representations. Such a representation is said t o beirreducible . It is not easy to tell by just looking at a representation whethe r is is reducible or not. We need to develop some tools. We begin with a more powe rful definition of irreducibilty. We first introduce the notion of an invariant subspace . Suppose we have a set{Aα}of linear maps acting on a vector space V. A subspace U⊆V is an invariant subspace for the set if x∈U⇒Aαx∈Ufor allAα. The set{Aα}isirreducible if the only invariant subspaces are Vitself and {0}. Conversely, if there is a non-trivial invariant subspace, then the set4of operators is reducible . If theAα’s posses a non-trivial invariant subspace U, and we decompose V=U⊕U/prime, whereU/primeis a complementary subspace, then, in a basis adapted to this decomposition, the matrices Aαtake the block-partitioned form of figure 5.1. 4Irreducibility is a property of the set as a whole. Any indivi dual matrix always has a non-trivial invariant subspace because it possesses at lea st one eigenvector. 5.2. REPRESENTATIONS 183                                                                                                                                         0AαU U Figure 5.1: Block partitioned reducible matrices. If we can find a5complementary subspace U/primewhich is also invariant, then we have the block partitioned form of figure 5.2.                                                                              00 AαU U Figure 5.2: Completely reducible matrices. We say that such matrices are completely reducible . When our linear op- erators are unitary with respect to some inner product, we ca n take the complementary subspace to be the orthogonal complement . This, by unitar- ity, is automatically be invariant. Thus, unitarity and red ucibility implies complete reducibility. Schur’s Lemma The most useful results concerning irreducibility come fro m: Schur’s Lemma : Suppose we have two sets of linear operators Aα:U→U, andBα:V→V, that act irreducibly on their spaces, and an intertwining operator Λ :U→Vsuch that ΛAα=BαΛ, (5.8) for allα, then either a) Λ = 0, or 5Remember that complementary subspaces are not unique. 184 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS b) Λ is 1-1 and onto (and hence invertible), in which case UandVhave the same dimension and Aα= Λ−1BαΛ. The proof is straightforward: The relation (5.8 ) shows that Ker (Λ)⊆Uand Im(Λ)⊆Vare invariant subspaces for the sets {Aα}and{Bα}respectively. Consequently, either Λ = 0, or Ker (Λ) = {0}and Im(Λ) = V. In the latter case Λ is 1-1 and onto, and hence invertible. Corollary: If{Aα}acts irreducibly on an n-dimensional vector space, and there is an operator Λ such that ΛAα=AαΛ, (5.9) then either Λ = 0 or Λ = λI. To see this observe that (5.9) remains true if Λ is replaced by (Λ −xI). Now det (Λ−xI) is a polynomial in xof degree n, and, by the fundamental theorem of algebra, has at least one root,x=λ. Since its determinant is zero, (Λ −λI) is not invertible, and so must vanish by Schur’s lemma. 5.2.2 Characters and Orthogonality Unitary Representations of Finite Groups LetGbe a finite group and let g/mapsto→D(g) be a representation of Gby matrices acting on a vector space V. Let ( x,y) denote a positive-definite, conjugate- symmetric, sesquilinear inner product of two vectors in V. From (,) we construct a new inner product /angbracketleft,/angbracketrightby averaging over the group /angbracketleftx,y/angbracketright=/summationdisplay g∈G(D(g)x,D(g)y). (5.10) It is easy to see that this new inner product remains positive definite, and in addition has the property that /angbracketleftD(g)x,D(g)y/angbracketright=/angbracketleftx,y/angbracketright. (5.11) This means that the maps D(g) :V→Vare unitary with respect to the new product. If we change basis to one that is orthonormal wit h respect to this new product then the D(g) become unitary matrices, with D(g−1) = D−1(g) =D†(g), whereD† ij(g) =D∗ ji(g) denotes the conjugate-transposed matrix. 5.2. REPRESENTATIONS 185 We conclude that representations of finite groups can always be taken to be unitary. This leads to the important consequence that f or such rep- resentations reducibility implies complete reducibility .Warning : In this construction it is essential that the sum over the g∈Gconverge. This is guaranteed for a finite group, but may not work for infinite gro ups. In par- ticular, non-compact Lie groups, such as the Lorentz group, have no finite dimensional unitary representations. Orthogonality of the Matrix Elements Now letDJ(g) :VJ→VJbe the matrices of an irreducible representation orirrep. HereJis a label which distinguishes inequivalent irreps from one another. We will use the symbol dim Jto denote the dimension of the rep- resentation vector space VJ. LetDKbe an irrep that is either identical to DJor inequivalent, and let Mijbe a matrix possessing the appropriate number of rows and col umns for productDJMDKto be defined, but otherwise arbitrary. The sum Λ =/summationdisplay g∈GDJ(g−1)MDK(g) (5.12) obeysDJ(g)Λ = ΛDK(g) for anyg. Consequently, Schur’s lemma tells us that Λil=/summationdisplay g∈GDJ ij(g−1)MjkDK kl(g) =λ(M)δilδJK. (5.13) We have written λ(M) to stress that the number λdepends on the chosen matrixM. Now take Mto be zero everywhere except for one entry of unity in rowjcolumnk. Then we have /summationdisplay g∈GDJ ij(g−1)DK kl(g) =λjkδil,δJK(5.14) where we have relabelled λto indicate its dependence on the location ( j,k) of the non-zero entry in M. We can find the constants λjkby assuming that K=J, settingi=l, and summing over i. We find |G|δjk=λjkdimJ. (5.15) Putting these results together we find that 1 |G|/summationdisplay g∈GDJ ij(g−1)DK kl(g) = (dimJ)−1δjkδilδJK. (5.16) 186 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS When our matrices D(g) are unitary, we can write this as 1 |G|/summationdisplay g∈G/parenleftbig DJ ij(g)/parenrightbig∗DK kl(g) = (dimJ)−1δikδjlδJK. (5.17) If we consider complex-valued functions G→Cas forming a vector space, then theDJ ijare elements of this space and are mutually orthogonal with respect to its natural inner product. There can be no more orthogonal functions on Gthan the dimension of the function space itself, which is |G|. We therefore have a constraint /summationdisplay J(dimJ)2≤|G| (5.18) that places a limit on how many inequivalent representation s can exist. In fact, as you will show later, the equality holds: the sum of th e squares of the dimensions of the inequivalent irreducible representatio ns is equal to the or- der ofG, and consequently the matrix elements form a complete ortho normal set of functions on G. Class functions and characters Because tr (C−1DC) = trD, (5.19) the trace of a representation matrix is the same for equivale nt representations. Further, because trD(g−1 1gg1) = tr/parenleftbig D−1(g1)D(g)D(g1)/parenrightbig = trD(g), (5.20) the trace is the same for all group elements in a conjugacy cla ss. The char- acter, χ(g)def= trD(g), (5.21) is therefore said to be a class function . By taking the trace of the matrix-element orthogonality rel ation we see that the characters χJ= trDJof the irreducible representations obey 1 |G|/summationdisplay g∈G/parenleftbig χJ(g)/parenrightbig∗χK(g) =1 |G|/summationdisplay idi/parenleftbig χJ i/parenrightbig∗χK i=δJK, (5.22) 5.2. REPRESENTATIONS 187 wherediis the number of elements in the i-th conjugacy class. The completeness of the matrix elements as functions on Gimplies that the characters form a complete orthogonal set of functions o n the space of conjugacy classes equipped with inner product /angbracketleftχ1,χ2/angbracketrightdef=1 |G|/summationdisplay idi/parenleftbig χ1 i/parenrightbig∗χ2 i. (5.23) Conseqently there are exactly as many inequivalent irreduc ible representa- tions as there are conjugacy classes in the group. Given a reducible representation, D(g), we can find out exactly which irrepsJit contains, and how many times, nJ, they occur. We do this forming thecompound character χ(g) = trD(g) (5.24) and observing that if we can find a basis in which D(g) = (D1(g)⊕D1(g)⊕···)/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright n1terms⊕(D2(g)⊕D2(g)⊕···)/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright n2terms⊕···,(5.25) then χ(g) =n1χ1(g) +n2χ2(g) +··· (5.26) From this we find nJ=/angbracketleftχ,χJ/angbracketright=1 |G|/summationdisplay idi(χi)∗χJ i. (5.27) There are extensive tables of group characters. Table 5.2 sh ows, for ex- ample, the characters of the group S4of permutations on 4 objects. Typical element and class size S4 (1) (12) (123) (1234) (12)(34) Irrep 1 6 8 6 3 A1 1 1 1 1 1 A2 1 -1 1 -1 1 E 2 0 -1 0 2 T1 3 1 0 -1 -1 T2 3 -1 0 1 -1 Table 5.2: Character table of S4 188 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS SinceχJ(e) = dimJwe see that the irreps A1andA2are one dimensional, thatEis two dimensional, and that T1,2are both three dimensional. Also we confirm that the sum of the squares of the dimensions 1 + 1 + 22+ 32+ 32= 24 = 4! is equal to the order of the group. As a further illustration of how to read table 5.2, let us veri fy the or- thonormality of the characters of the representations T1andT2. We have /angbracketleftχT1,χT2/angbracketright=1 |G|/summationdisplay idi/parenleftbig χT1 i/parenrightbig∗χT2 i=1 24[1·3·3−6·1·1+8·0·0−6·1·1+3·1·1] = 0, while /angbracketleftχT1,χT1/angbracketright=1 |G|/summationdisplay idi/parenleftbig χT1 i/parenrightbig∗χT1 i=1 24[1·3·3+6·1·1+8·0·0+6·1·1+3·1·1] = 1. The sum giving/angbracketleftχT2,χT2/angbracketright= 1 is identical to this. Exercise 5.14 : LetD1andD2be representations with characters χ1(g) and χ2(g) respectively. Show that the character of the direct produc t representa- tionD1⊗D2is given by χ1⊗2(g) =χ1(g)χ2(g). 5.2.3 The Group Algebra Given a finite group G, we construct a vector space C(G) whose basis vectors are in one-to-one correspondence with the elements of the gr oup. We denote the vector corresponding to the group element gby the boldface symbol g. A general element of C(G) is therefore a formal sum x=x1g1+x2g2+···+x|G|g|G|. (5.28) We take products of these sums by using the group multiplicat ion rule. If g1g2=g3we set g1g2=g3, and require the product to be distributive with respect to vector-space addition. Thus gx=x1gg1+x2gg2+···+x|G|gg|G|. (5.29) 5.2. REPRESENTATIONS 189 The resulting mathematical structure is called the group algebra . It was introduced by Frobenius. The group algebra, considered as a vector space, is automati cally a rep- resentation. We define the natural action of GonC(G) by setting D(g)gi=ggi=gjDji(g). (5.30) The matrices Dji(g) make up the regular representation. Because the list gg1,gg2,...is a permutation of the list g1,g2,..., their entries consist of 1’s and 0’s, with exactly one non-zero entry in each row and each c olumn. Exercise 5.15 : Show that the character of the regular representation has χ(e) = |G|, andχ(g) = 0, forg/negationslash=e. Exercise 5.16 : Use the previous exercise to show that the number of times anndimensional irrep occurs in the regular representation is n. Deduce that |G|=/summationtext J(dimJ)2, and from this construct the completeness proof for the representations and characters. Projection Operators A representation DJof the group automatically provides a representation of the group algebra. We simply set DJ(x1g1+x2g2+···)def=x1DJ(g1) +x2DJ(g2) +···. (5.31) Certain linear combinations of group elements turn out to be very useful because the corresponding matrices can be used to project ou t vectors with desirable symmetry properties. Consider the elements eJ αβ=dimJ |G|/summationdisplay g∈G/bracketleftbig DJ αβ(g)/bracketrightbig∗g (5.32) of the group algebra. These have the property that g1eJ αβ=dimJ |G|/summationdisplay g∈G/bracketleftbig DJ αβ(g)/bracketrightbig∗(g1g) =dimJ |G|/summationdisplay g∈G/bracketleftbig DJ αβ(g−1 1g)/bracketrightbig∗g 190 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS =/bracketleftbig DJ αγ(g−1 1)/bracketrightbig∗dimJ |G|/summationdisplay g∈G/bracketleftbig DJ γβ(g)/bracketrightbig∗g =eJ γβDJ γα(g1). (5.33) In going from the first to the second line we have changed summa tion vari- ables from g→g−1 1g, and going from the second to the third line we have used the representation property to write DJ(g−1 1g) =DJ(g−1 1)DJ(g). Fromg1eJ αβ=eJ γβDJ γα(g1) and the matrix-element orthogonality, it fol- lows that eJ αβeK γδ=dimJ |G|/summationdisplay g∈G/bracketleftbig DJ αβ(g)/bracketrightbig∗geK γδ =dimJ |G|/summationdisplay g∈G/bracketleftbig DJ αβ(g)/bracketrightbig∗DK /epsilon1γ(g)eK /epsilon1δ =δJKδα/epsilon1δβγeK /epsilon1δ =δJKδβγeJ αδ. (5.34) For eachJ, this multiplication rule of the eJ αβis identical to that of matrices having zero entries everywhere except for the ( α,β)-th, which is a “1.” There are (dimJ)2of these eJ αβfor eachn-dimensional representation J, and they are linearly independent. Because/summationtext J(dimJ)2=|G|, they form a basis for the algebra. In particular every element of Gcan be reconstructed as g=/summationdisplay JDJ ij(g)eJ ij. (5.35) We can also define the useful objects PJ=/summationdisplay ieJ ii=dimJ |G|/summationdisplay g∈G/bracketleftbig χJ(g)/bracketrightbig∗g. (5.36) They have the property PJPK=δJKPK,/summationdisplay JPJ=I, (5.37) where Iis the identity element of C(G). The PJare therefore projection operators composing a resolution of the identity. Their uti lity resides in the fact that when D(g) is a reducible representation acting on a linear space V=/circleplusdisplay JVJ, (5.38) 5.2. REPRESENTATIONS 191 then setting g→D(g) in the formula for PJresults in a projection matrix fromVonto the irreducible component VJ. To see how this comes about, let v∈Vand, for any fixed p, set vi=eJ ipv, (5.39) where eJ ipvshould be understood as shorthand for D(eJ ip)v. Then D(g)vi=geJ ipv=eJ jpvDJ ji(g) =vjDJ ji(g). (5.40) We see the vi, if not all zero, are basis vectors for VJ. Since PJis a sum of theeJ ij, the vector PJvis a sum of such vectors, and therefore lies in VJ. The advantage of using PJover any individual eJ ipis that PJcan be computed from character table, i.e.its construction does not require knowledge of the irreducible representation matrices. The algebra of classes If a conjugacy class Ciconsists of the elements {g1,g2,...gdi}, we can define Cito be the corresponding element of the group algebra: Ci=1 di(g1+g2+···gdi). (5.41) (The factor of 1 /diis a conventional normalization.) Because conjugation merely permutes the elements of a conjugacy class, we have g−1Cig=Ci for all g∈C(G). The Citherefore commute with every element of C(G). Conversely any element of C(G) that commutes with everything in C(G) must be a linear combination C=c1C1+c2C2+.... The subspace of C(G) consisting of sums of the classes is therefore the centreZ[C(G)] of the group algebra. Because the product CiCjcommutes with everything, it lies in Z[C(G)] and so there are constants cijksuch that CiCj=/summationdisplay kcijkCk. (5.42) We can regard the Cias being linear maps from Z[C(G)] to itself, whose associated matrices have entries ( Ci)k j=cijk. These matrices commute, and can be simultaneously diagonalized. We will leave it as e xercise for the reader to demonstrate that CiPJ=/parenleftbiggχJ i χJ 0/parenrightbigg PJ. (5.43) 192 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS HereχJ 0≡χJ {e}= dimJ. The common eigenvectors of the Ciare therefore the projection operators PJ, and the eigenvalues λJ i=χJ i/χJ 0are, up to nor- malization, the characters. Equation (5.43) provides a con venient method for computing the characters from knowledge only of the coeffi cientscijk appearing in the class multiplication table. Once we have fo und the eigen- valuesλJ i, we recover the χJ iby noting that χJ 0is real and positive, and that/summationtext idi|χJ i|2=|G|. Exercise 5.17 : Use Schur’s lemma to show that for an irrep DJ(g) we have 1 di/summationdisplay g∈CiDJ jk(g) =1 dimJδjkχJ i, and hence establish (5.43). 5.3 Physics Applications 5.3.1 Quantum Mechanics When a group G={gi}acts on a mechanical system, then Gwill act as set of linear operators D(g) on the Hilbert space Hof the corresponding quantum system. ThusHwill be a representation6space forG. If the group is a symmetry of the system then the D(g) will commute with the hamiltonian ˆH. If this is so, and if we can decompose H=/circleplusdisplay irrepsJHJ (5.44) intoˆH-invariant irreps of Gthen Schur’s lemma tells us that in each HJthe hamiltonian ˆHwill act as a multiple of the identity operator. In other word s every state inHJwill be an eigenstate of ˆHwith a common energy EJ. This fact can greatly simplify the task of finding the energy l evels. If an irrepJoccurs only once in the decomposition of Hthen we can find the eigenstates directly by applying the projection operator PJto vectors inH. 6The rules of quantum mechanics only require that D(g1)D(g2) =eiφ(g1,g2)D(g1g2). A set of matrices that obeys the group multiplication rule “u p to a phase” is called a projective (orray) representation. In many cases, however, we can choose the D(g) so thatφis not needed. This is the case in all the examples we discuss. 5.3. PHYSICS APPLICATIONS 193 If the irrep occurs nJtimes in the decomposition, then PJwill project to the reducible subspace HJ⊕HJ⊕···HJ/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright nJcopies=M⊗HJ. HereMis annJdimensional multiplicity space . The hamiltonian ˆHwill act inMas annJ-by-nJmatrix. In other words, if the vectors |n,i/angbracketright≡|n/angbracketright⊗|i/angbracketright∈M⊗H J (5.45) form a basiorM⊗HJ, withnlabelling which copy of HJthe vector|n,i/angbracketright lies in, then ˆH|n,i/angbracketright=|m,i/angbracketrightHJ mn, D(g)|n,i/angbracketright=|n,j/angbracketrightDJ ji(g). (5.46) Diagonalizing HJ nmprovides us with njˆH-invariant copies of HJand gives us the energy eigenstates. Consider, for example, the molecule C 60(buckminsterfullerine) consisting of 60 carbon atoms in the form of a soccer ball. The chemically active electrons can be treated in a tight-binding approximation i n which the Hilbert space has dimension 60 — one π-orbital basis state for each each carbon atom. The geometric symmetry group of the molecule is Yh=Y×Z2, whereYis the rotational symmetry group of the icosohedron (a subgrou p of SO(3)) and Z2is the parity inversion σ:r/mapsto→−r. The characters of Yare displayed in table 5.3. Typical element and class size Y eC5C2 5C2C3 Irrep 1 12 12 15 20 A 1 1 1 1 1 T1 3τ−1−τ-1 0 T2 3−τ τ−1-1 0 G 4 -1 -1 0 1 H 5 0 0 1 -1 Table 5.3: Character table for the group Y. 194 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS In this table τ=1 2(√ 5−1) denotes the golden mean. The class C5is the set of 2π/5 rotations about an axis through the centres of a pair of anti podal pentagonal faces, the class C3is the set of of 2 π/3 rotations about an axis through the centres of a pair of antipodal hexagonal faces, a ndC2is the set ofπrotatations through the midpoints of a pair of antipodal edg es, each lying between two adjacent hexagonal faces. 123 0g h h ggL=4tgu g hut2u hg t t1g 1u u g gut hg t1u agL=3 L=2 L=1 L=02ug2 −3−2−1E 60C Figure 5.3: A sketch of the tight-binding electronic energy levels of C 60. The geometric symmetry group acts on the 60-dimensional Hil bert space by permuting the basis states concurrently with their associa ted atoms. Figure 5.3 shows how the 60 states are disposed into energy levels.7Each level is labelled by a lower case letter specifying the irrep of Y, and by a subscript gorustanding for gerade (German for even) orungerade (German for odd) that indicates whether the wavefunction is even or odd under the inversion σ:r/mapsto→−r. The buckyball is roughly spherical, and the lowest 25 states can be thought as being derived from the angular-momentum eigenst ates withL= 0,1,2,3,4,that classify the energy levels for an electron moving on a pe rfect sphere. In the many-electron ground-state, the 30 single-p article states with energy below E <0 are each occupied by pairs of spin up/down electrons. The 30 states with E >0 are empty. 7After R. C. Haddon, L. E. Brus, K. Raghavachari, Chem. Phys. Lett. 125(1986) 459. 5.3. PHYSICS APPLICATIONS 195 To explain, for example, why three copies of T1appear, and why two of these are T1uand oneT1g, we must investigate the manner in which the 60-dimensional Hilbert space decomposes into irreducible representations of 120-element group Yh. Problem 5.23 leads us through this computation, and shows that no irrep of Yhoccurs more that three times. In finding the energy levels, we therefore never have to diagonalize a bigger than 3-by-3 matrix. The equality of the energies of the hgandgglevels atE=−1 is an accidental degeneracy . It is not required by the symmetry, and will presum- ably disappear in a more sophisticated calculation. The app earance of many “accidental” degeneracies in an energy spectrum hints that there may be a hidden symmetry that arises from something beyond geometry. For example, in the Schr¨ odinger spectrum of the hydrogen atom all states with the same principal quantum number nhave the same energy although they correspond to different irreps L= 1,...,n−1 of O(3). This degeneracy occurs because the classical Kepler-orbit problem has symmetry group O(4) , rather than the na¨ ıvely expected O(3) rotational symmetry. 5.3.2 Vibrational spectrum of H 2O The small vibrations of a mechanical system with ndegrees of freedom are governed by a Lagrangian of the form L=1 2˙xTM˙x−1 2xTVx (5.47) whereMandVare symmetric n-by-nmatrices, and with Mbeing positive definite. This Lagrangian leads to the equations of motion M¨x=Vx (5.48) We look for normal mode solutions x(t)∝eiωitxi, where the vectors xiobey −ω2 iMxi=Vxi. (5.49) The normal-mode frequencies are solutions of the secular eq uation det (V−ω2M) = 0, (5.50) and modes with distinct frequencies are orthogonal with res pect to the inner product defined by M, /angbracketleftx,y/angbracketright=xTMy. (5.51) 196 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS We are interested in solving this problem for vibrations abo ut the equi- librium configuration of a molecule. Suppose this equilibri um configuration has a symmetry group G. This gives rise to an n-dimensional representation on the space of x’s in which g:x/mapsto→D(g)x, (5.52) leaves both the intertia matrix Mand the potential matrix Vunchanged. [D(g)]TMD(g) =M, [D(g)]TVD(g) =V. (5.53) Consequently, if we have an eigenvector xiwith frequency ωi, −ω2 iMxi=Vxi (5.54) we see that D(g)xialso satisfies this equation. The frequency eigenspaces are therefore left invariant by the action of D(g), and barring accidental degeneracy, there will be a one-to-one correspondence betw een the frequency eigenspaces and the irreducible representations occurrin g inD(g). Consider, for example, the vibrational modes of the water mo lecule H 2O. This familiar molecule has symmetry group C2vwhich is generated by two elements: a rotation athroughπabout an axis through the oxygen atom, and a reflection bin the plane through the oxygen atom and bisecting the angle between the two hydrogens. The product abis a reflection in the plane defined by the equilibrium position of the three atoms. The re lations are a2=b2= (ab)2=e, and the characters are displayed in table 5.4. class and size C2ve a b ab Irrep 1 1 1 1 A1 1 1 1 1 A2 1 1 -1 -1 B1 1 -1 1 -1 B2 1 -1 -1 1 Table 5.4: Character table of C2v. The group C2vis Abelian, so all the representations are one dimensional. 5.3. PHYSICS APPLICATIONS 197 To find out what representations occur when C2vacts, we need to find the character of its action D(g) on the nine-dimensional vector x= (xO,yO,zO,xH1,yH1,zH1,xH2,yH2,zH2). (5.55) Here the coordinates xH2,yH2,zH2etc.denote the displacements of the la- belled atom from its equilibrium position. We take the molecule as lying in the xyplane, with the zpointing towards us. H H1 2O H H2 2y xxy xyH H1 1OO Figure 5.4: Water Molecule. The effect of the symmetry operations on the atomic displacem ents is D(a)x= (−xO,+yO,−zO,−xH2,+yH2,−zH2,−xH1,+yH1,−zH1) D(b)x= (−xO,+yO,+zO,−xH2,+yH2,+zH2,−xH1,+yH1,+zH1) D(ab)x= (+xO,+yO,−zO,+xH1,+yH1,−zH1,+xH2,+yH2,−zH2). Notice how the transformations D(a),D(b) have interchanged the displace- ment co-ordinates of the two hydrogen atoms. In calculating the character of a transformation we need look only at the effect on atoms tha t are left fixed — those that are moved have matrix elements only in non-d iagonal positions. Thus, when computing the compound characters fo ra b, we can focus on the oxygen atom. For abwe need to look at all three atoms. We find χD(e) = 9, χD(a) =−1 + 1−1 =−1, χD(b) =−1 + 1 + 1 = 1 , χD(ab) = 1 + 1−1 + 1 + 1−1 + 1 + 1−1 = 3. 198 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS By using the orthogonality relations, we find the decomposit ion  9 −1 1 3 = 3 1 1 1 1 + 1 1 −1 −1 + 2 1 −1 1 −1 + 3 1 −1 −1 1 (5.56) or χD= 3χA1+χA2+ 2χB1+ 3χB2. (5.57) Thus, the nine-dimensional representation decomposes as D= 3A1⊕A2⊕2B1⊕3B2. (5.58) How do we exploit this? First we cut out the junk. Out of the nin e modes, six correspond to easily identified zero-frequency m otions – three of translation and three rotations. A translation in the xdirection would have xO=xH1=xH2=ξ, all other entries being zero. This displacement vector changes sign under both aandb, but is left fixed by ab. This behaviour is characteristic of the representation B2. Similarly we can identify A1as translation in y, andB1as translation in z. A rotation about the yaxis makeszH1=−zH2=φ. This is left fixed by a, but changes sign under band ab, so theyrotation mode is A2. Similarly, rotations about the xandzaxes correspond to B1andB2respectively. All that is left for genuine vibrational modes is 2A1⊕B2. We now apply the projection operator PA1=1 4[(χA1(e))∗D(e) + (χA1(a))∗D(b) + (χA1(b))∗D(b) + (χA1(ab))∗D(ab)] (5.59) tovH1,x, a small displacement of H1in the x direction. We find PA1vH1,x=1 4(vH1,x−vH2,x−vH2,x+vH1,x) =1 2(vH1,x−vH2,x). (5.60) This mode is an eigenvector for the vibration problem. If we apply PA1tovH1,yandvO,ywe find PA1vH1,y=1 2(vH1,y+vH2,y), PA1vO,y=vO,y, (5.61) 5.3. PHYSICS APPLICATIONS 199 but we are not quite done. These modes are contaminated by the ytrans- lation direction zero mode, which is also in an A1representation. After we make our modes orthogonal to this, there is only one left, a nd this has yH1=yH2=−yOmO/(2mH) =a1, all other components vanishing. We can similarly find vectors corresponding to B2as PB2vH1,x=1 2(vH1,x+vH2,x) PB2vH1,y=1 2(vH1,y−vH2,y) PB2vO,x=vO,x and these need to be cleared of both translations in the xdirection and rotations about the zaxis, both of which transform under B2. Again there is only one mode left and it is yH1=−yH2=αxH1=αxH2=βx0=a2 (5.62) whereαis chosen to ensure that there is no angular momentum about O, andβto make the total xlinear momentum vanish. We have therefore found three true vibration eigenmodes, two transforming un derA1and one underB2as advertised earlier. The eigenfrequencies, of course, de pend on the details of the spring constants, but now that we have the e igenvectors we can just plug them in to find these. 5.3.3 Crystal Field Splittings A quantum mechanical system has a symmetry Gif the hamiltonian ˆHobeys D−1(g)ˆHD(g) =ˆH, (5.63) for some group action D(g) :H→H on the Hilbert space. If follows that the eigenspaces,Hλ, of states with a common eigenvalue, λ, are invariant subspaces for the representation D(g). We often need to understand how a degeneracy is lifted by pert urbations that break Gdown to a smaller subgroup H. Ann-dimensional irreducible representation of Gis automatically a representation of any subgroup of G, but in general it is no longer be irreducible. Thus the n-fold degenerate level is split into multiplets, one for each of the irreducib le representations 200 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS ofHcontained in the original representation. The manner in whi ch an orig- inally irreducible representation decomposes under restr iction to a subgroup is known as the branching rule for the representation. A physically important case is given by the breaking of the fu ll SO(3) rotation symmetry of an isolated atomic hamiltonian by a cry stal field Sup- pose the crystal has octohedral symmetry. The characters of the octohedral group are displayed in table 5.5. Class(size) Oe C 3(8)C2 4(3)C2(6)C4(6) A11 1 1 1 1 A21 1 1 -1 -1 E 2 -1 2 0 0 F23 0 -1 1 -1 F13 0 -1 -1 1 Table 5.5: Character table of the octohedral group O. The classes are lableled by the rotation angles, C2being a twofold rotation axis (θ=π),C3a threefold axis ( θ= 2π/3),etc.. The chacter of the J=lrepresentation of SO(3) is χl(θ) =sin(2l+ 1)θ/2 sinθ/2, (5.64) and the first few χl’s evaluated on the rotation angles of the classes of Oare dsiplayed in table 5.6. Class(size) le C 3(8)C2 4(3)C2(6)C4(6) 01 1 1 1 1 13 0 -1 -1 -1 25 -1 1 1 -1 37 1 -1 -1 -1 49 0 1 1 1 Table 5.6: Characters evaluated on rotation classes 5.4. FURTHER EXERCISES AND PROBLEMS 201 The 9-fold degenerate l= 4 multiplet therefore decomposes as  9 0 1 1 1 = 1 1 1 1 1 + 2 −1 2 0 0 + 3 0 −1 −1 1 + 3 0 −1 1 −1 , (5.65) or χ4 SO(3)=χA1+χE+χF1+χF2. (5.66) The octohedral crystal field splits the nine states into four multiplets with symmetries A1,E,F1,F2and degeneracies 1, 2, 3 and 3, respectively. We have considered only the simplest case here, ignoring the complica- tions introduced by reflection symmetries, and by 2-valued s pinor represen- tations of the rotation group. 5.4 Further Exercises and Problems We begin with some technologically important applications of group theory to cryptography and number theory. Exercise 5.18 : The set Znforms a group under multiplication only when nis a prime number. Show, however, that the subset U( Zn)⊂Znof elements of Znthat are co-prime to nis a group. It is the group of units of the ring Zn. Exercise 5.19 :Cyclic groups . A group Gis said to be cyclic if its elements consist of powers anof of an element a, called the generator . The group will be of finite order |G|=mifam=a0=efor somem∈Z+. a) Show that a group of prime order is necessarily cyclic, and that any element other than the identity can serve as its generator. ( Hint: Let abe any element other than eand consider the subgroup consisting of powersam.) b) Show that any subgroup of a cyclic group is itself cyclic. Exercise 5.20 :Cyclic groups and cryptography . In a large cyclic group G it can be relatively easy to compute ax, but to recover xgivenh=axone might have to compute ayand compare it with hfor every 1 < y <|G|. If |G|has several hundred digits, such a brute force search could t ake longer than the age of the universe. Rather more efficient algorithms for this discrete logarithm problem exist, but the difficulty is still sufficient for it to be useful in cryptopgraphy. 202 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS a)Diffie-Hellman key exchange . This algorithm allows Alice and Bob to establish a secret key that can be used with a conventional cy pher with- out Eve, who is listening to their conversation, being able t o reconstruct it. Alice choses a random element g∈Gand an integer xbetween 1 and |G|and computes gx. She sends gandgxto Bob, but keeps xto herself. Bob chooses an integer yand computes gyandgxy= (gx)y. He keeps ysecret and sends gyto Alice, who computes gxy= (gy)x. Show that, although Eve knows g,gyandgx, she cannot obtain Alice and Bob’s secret keygxywithout solving the discrete logarithm problem. b)ElGamal public key encryption . This algorithm, based on Diffie-Hellman, was invented by the Egyptian cryptographer Taher Elgamal. I t is a component of PGP and and other modern encryption packages. T o use it, Alice first chooses a random integer xin the range 1 to |G|and computesh=ax. She publishes a description of G, together with the elementshanda, as her public key. She keeps the integer xsecret. To send a message mto Alice, Bob chooses an integer yin the same range and computes c1=ay,c2=mhy. He transmits c1andc2to Alice, but keepsysecret. Alice can recover mfromc1,c2by computing c2(cx 1)−1. Show that, although Eve knows Alice’s public key and has over heardc1 andc2, she nonetheless cannot decrypt the message without solvin g the discrete logarithm problem. Popular choices for Gare subgroups of ( Zp)×, for large prime p. (Zp)×is itself cyclic (can you prove this?), but is unsuitable for technica l reasons. Exercise 5.21 :Modular arithmetic and number theory . An integer ais said to be a quadratic residue modpif there is an rsuch thata=r2(modp). Letpbe an odd prime. Show that if r2 1=r2 2(modp) thenr1=±r2(modp), and thatr/negationslash=−r(modp). Deduce that exactly one half of thep−1 non-zero elements of Zpare quadratic residues. Now consider the Legendre symbol /parenleftbigga p/parenrightbigg def=  0, a = 0, 1, a a quadratic residue (mod p), −1anot a quadratic residue (mod p). Show that /parenleftbigga p/parenrightbigg/parenleftbiggb p/parenrightbigg =/parenleftbiggab p/parenrightbigg , and so the Legendre symbol forms a one-dimensional represen tation of the multiplicative group ( Zp)×. Combine this fact with the character orthogonality 5.4. FURTHER EXERCISES AND PROBLEMS 203 theorem to give an alternative proof that precisely half the p−1 elements of (Zp)×are quadratic residues. (Hint: To show that the product of tw o non- residues is a residue, observe that the set of residues is a no rmal subgroup of (Zp)×, and consider the multiplication table of the resulting quo tient group.) Exercise 5.22 :More practice with modular arithmetic . Again let pbe an odd prime. Prove Euler’s theorem that a(p−1)/2(modp) =/parenleftbigga p/parenrightbigg . (Hint: Begin by showing that the usual school-algebra proof that an equa- tion of degree ncan have no more than nsolutions remains valid for arith- metic modulo a prime number, and so a(p−1)/2= 1 (modp) can have no more than(p−1)/2 roots. Cite Fermat’s little theorem to show that these root s must be the quadratic residues. Cite Fermat again to show tha t the quadratic non-residues must then have a(p−1)/2=−1 (modp).) The harder-to-prove law of quadratic reciprocity asserts that for p,qodd primes, we have (−1)(p−1)(q−1)/4/parenleftbiggp q/parenrightbigg =/parenleftbiggq p/parenrightbigg . Problem 5.23 :Buckyball spectrum. Consider the symmetry group of the C 60 buckyball molecule of figure 5.3. a) Starting from the character table of the orientation-pre serving icosohe- dral groupY(table 5.3), and using the fact that the Z2parity inversion σ:r→−rcombines with g∈Yso thatDJg(σg) =DJg(g), whilst DJu(σg) =−DJu(g), write down the character table of the extended groupYh=Y×Z2that acts as a symmetry on the C 60molecule. There are now ten conjugacy classes, and the ten representations w ill be la- belledAg,Au,etc. Verify that your character table has the expected row-orthogonality properties. b) By counting the number of atoms left fixed by each group oper ation, compute the compound character of the action of Yhon the C 60molecule. (Hint: Examine the pattern of panels on a regulation soccer b all, and deduce that four carbon atoms are left unmoved by operations in the classσC2.) c) Use your compound character from part b), to show that the 6 0-dimensional Hillbert space decomposes as HC60=Ag⊕T1g⊕2T1u⊕T2g⊕2T2u⊕2Gg⊕2Gu⊕3Hg⊕2Hu, consistent with the energy-levels sketched in figure 5.3. 204 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS Problem 5.24 :The Frobenius-Schur Indicator. Recall that a real or pseudo- real representation is one such that D(g)∼D∗(g), and for unitary matrices D we haveD∗(g) = [DT(g)]−1. In this unitary case D(g) being real or pseudo- real is equivalent to the statement that there exists an inve rtible matrix F such that FD(g)F−1= [DT(g)]−1. We can rewrite this statement as DT(g)FD(g) =F, and soFcan be inter- preted as the matrix representing a G-invariant quadratic form. i) Use Schur’s lemma to show that when Dis irreducible the matrix Fis unique up to an overall constant. In other words, DT(g)F1D(g) =F1 andDT(g)F2D(g) =F2for allg∈Gimplies that F2=λF1. Deduce that for irreducible Dwe haveFT=±F. ii) By reducing Fto a suitable canonical form, show that Fis symmetric (F=FT) in the case that D(g) is a real representation, and Fis skew symmetric ( F=−FT) whenD(g) is a pseudo-real representation. iii) Now let Gbe afinite group. For any matrix U, the sum FU=1 |G|/summationdisplay g∈GDT(g)UD(g) is aG-invariant matrix. Deduce that FUis always zero when D(g) is neither real nor pseudo-real, and, by specializing both Uand the indices onFU, show that in the real or pseudo-real case /summationdisplay g∈Gχ(g2) =±/summationdisplay g∈Gχ(g)χ(g), whereχ(g) = trD(g) is the character of the irreducible representation D(g). Deduce that the Frobenius-Schur indicator κdef=1 |G|/summationdisplay g∈Gχ(g2) takes the value +1, −1, or 0 when D(g) is, respectively, real, pseudo-real, or not real. iv) Show that the identity representation occurs in the deco mposition of the tensor product D(g)⊗D(g) of an irrep with itself if, and only if, D(g) is real or pseudo-real. Given a basis eifor the vector space Von which D(g) acts, show the matrix Fcan be used to construct the basis for the identity-representation subspace Vidin the decomposition V⊗V=/circleplusdisplay irrepsJVJ. 5.4. FURTHER EXERCISES AND PROBLEMS 205 Problem 5.25 :Induced Representations . Suppose we know a representation DW(h) :W→Wfor a subgroup H⊂G. From this representation we can construct an induced representation IndG H(DW) for the larger group G. The construction cleverly combines the coset space G/H with the representation spaceWto make a (usually reducible) representation space IndG H(W) of di- mension|G/H|×dimW. Recall that there is a natural action of Gon the coset space G/H. Ifx= {g1,g2,...}∈G/H thengxis the coset{gg1,gg2,...}.We select from each cosetx∈G/Ha representative element ax, and observe that the product gax can be decomposed as gax=agxh, whereagxis the selected representative from the coset gxandhis some element of H. Next we introduce a basis |n,x/angbracketrightfor IndG H(W). We use the symbol “0” to label the coset {e}, and take |n,0/angbracketrightto be the basis vectors for W. Forh∈Hwe can therefore set D(h)|n,0/angbracketrightdef=|m,0/angbracketrightDW mn(h). We also define the result of the action of axon|n,0/angbracketrightto be the vector|n,x/angbracketright: D(ax)|n,0/angbracketrightdef=|n,x/angbracketright. We may now obtain the the action of a general element of Gon the vectors |n,x/angbracketrightby requiring D(g) to be representation, and so computing D(g)|n,x/angbracketright=D(g)D(ax)|n,0/angbracketright =D(gax)|n,0/angbracketright =D(agxh)|n,0/angbracketright =D(agx)D(h)|n,0/angbracketright =D(agx)|m,0/angbracketrightDW mn(h) =|m,gx/angbracketrightDW mn(h). i) Confirm that the action D(g)|n,x/angbracketright=|m,gx/angbracketrightDW mn(h), withhobtained fromgandxviathe decomposition gax=agxh, does indeed define a representation of G. Show also that if we set |f/angbracketright=/summationtext n,xfn(x)|n,x/angbracketright, then the action of gon the components takes fn(x)/mapsto→DW nm(h)fm(g−1x). ii) Letf(h) be a class function on H. Let us extend it to a function on G by settingf(g) = 0 ifg /∈H, and define IndG H[f](s) =1 |H|/summationdisplay g∈Gf(g−1sg). 206 CHAPTER 5. GROUPS AND GROUP REPRESENTATIONS Show that IndG H[f](s) is a class function on G, and further show that if χWis the character of the starting representation for Hthen IndG H[χW] is the character of the induced representation of G. (Hint, only fixed points of the G-action onG/H contribute to the character, and gx=x means that gax=axh. ThusDW(h) =DW(a−1 xgax).) iii) Given a representation DV(g) :V→VofGwe can trivially obtain a (generally reducible) representation ResG H(V) ofH⊂Gby restricting G toH. Define the usual inner product on the group functions by /angbracketleftφ1,φ2/angbracketrightG=1 |G|/summationdisplay g∈Gφ1(g−1)φ2(g), and show that if ψis a class function on Handφa class function on G then /angbracketleftψ,ResG H[φ]/angbracketrightH=/angbracketleftIndG H[ψ],φ/angbracketrightG. Thus, IndG Hand ResG Hare, in some sense, adjoint operations. Mathe- maticians would call them a pair of mutually adjoint functors . iv) By applying the result from part (iii) to the characters o f the irreducible representations of GandH, deduce Frobenius’ reciprocity theorem : The number of times an irrep DJ(g) ofGoccurs in the representation induced from an irrep DK(h) ofHis equal to the number of times that DKoccurs in the decomposition of DJinto irreps of H. The representation of the Poincar´ e group (= the SO(1 ,3) Lorentz group to- gether with space-time translations) that classifies the st ates of a spin- Jele- mentary particle are those induced from the spin- Jrepresentation of its SO(3) rotation subgroup. The quantum state of a mass melementary particle is therefore of the form |k,σ/angbracketrightwherekis the particle’s four-momentum, which lies is the coset SO(1 ,3)/SO(3), and σis the label from the |J,σ/angbracketrightspin state. Chapter 6 Lie Groups Lie groups are named after the Norwegian mathematician Soph us Lie. They consist of a manifold Gequipped with a group multiplication rule ( g1,g2)/mapsto→g3 which is a smooth function of the g’s, as is the operation of taking the inverse of a group element. The most commonly met examples in physics are the infinite families of matrix groups GL(n), SL(n), O(n), SO(n), U(n), SU(n), and Sp(n), togther with the family of five exceptional Lie groups: G 2, F4, E6, E7, and E 8, which have applications in string theory. One of the properties of a Lie group is that, considered as a ma nifold, the neighbourhood of any point looks exactly like that of any other. The group’s dimension and most of its structure can be understoo d by examining the immediate vicinity any chosen point, which we may as well take to be the identity element. The vectors lying in the tangent space at the identity element make up the Lie algebra of the group. Computations in the Lie algebra are often easier than those in the group, and provide much of the same information. This chapter will be devoted to studying t he interplay between the Lie group itself and this Lie algebra of infinites imal elements. 6.1 Matrix Groups TheClassical Groups are described in a book with this title by Hermann Weyl. They are subgroups of the general linear group , GL(n,F), which con- sists of invertible n-by-nmatrices over the field F. We will mostly consider the cases F=CorF=R. A near-identity matrix in GL( n,R) can be written g=I+/epsilon1AwhereA 207 208 CHAPTER 6. LIE GROUPS is an arbitrary n-by-nreal matrix. This matrix contains n2real entries, so we can move away from the identity in n2distinct directions. The tangent space at the identity, and hence the group manifold itself, i s therefore n2 dimensional. The manifold of GL( n,C) hasn2complex dimensions, and this corresponds to 2 n2real dimensions. If we restrict the determinant of a GL( n,F) matrix to be unity, we get thespecial linear group , SL(n,F). An element near the identity in this group can still be written as g=I+/epsilon1A, but since det (I+/epsilon1A) = 1 +/epsilon1tr(A) +O(/epsilon12) (6.1) this requires tr( A) = 0. The restriction on the trace means that SL( n,R) has dimension n2−1. 6.1.1 The Unitary and Orthogonal Groups Perhaps the most important of the matrix groups are the unita ry and or- thogonal groups. The Unitary group The unitary group U( n) comprises the set of n-by-ncomplex matrices Usuch thatU†=U−1. If we consider matrices near the identity U=I+/epsilon1A, (6.2) with/epsilon1real, then unitarity requires I+O(/epsilon12) = (I+/epsilon1A)(I+/epsilon1A†) =I+/epsilon1(A+A†) +O(/epsilon12), (6.3) soAij=−A∗ jiandAis skew hermitian. A complex skew-hermitian matrix contains n+ 2×1 2n(n−1) =n2 real parameters. In this counting the first “ n” is the number of entries on the diagonal, each of which must be of the form itimes a real number. The n(n−1)/2 is the number of entries above the main diagonal, each of whi ch can be an arbitrary complex number. The number of real dimens ions in the 6.1. MATRIX GROUPS 209 group manifold is therefore n2. The rows or columns in the matrix Uform an orthonormal set of vectors. Their entries are therefore b ounded,|Uij|≤1, and this property leads to the n2dimensional group manifold of U( n) being a compact set. When a group manifold is compact, we say that the group itself is a compact group . There is a natural notion of volume on a group manifold and compact Lie groups have finite total volume. Because of th is, they have many properties in common with the finite groups we studied in the last chapter. Recall that a group is simple if it possesses no invariant subgroups. U( n) is not simple. Its centre is an invariant U(1) subgroup consi sting of matrices of the form U=eiθI. The special unitary group SU(n), consists of n-by-n unimodular (having determinant +1 ) unitary matrices. It is not strictly simple because its center Zconsists of the discrete subgroup of matrices Um=ωmIwithωann-th root of unity, and this is an invariant subgroup. BecauseZ, its only invariant subgroup, is not a continuous group, SU( n) is counted as being simple in Lie theory. With U=I+/epsilon1A, as above, the unimodularity imposes the additional constraint on Athat trA= 0, so the SU(n) group manifold is n2−1 dimensional. The Orthogonal Group The orthogonal group O( n), consists of the the set of real matrices Owith the property that OT=O−1. For a matrix in the neighbourhood of the identity,O=I+/epsilon1A, this condition requires that Abe skew symmetric: Aij=−Aij. Skew symmetric real matrices have n(n−1)/2 independent entries, and so the group manifold of O( n) isn(n−1)/2 dimensional. The conditionOTO=Imeans that the rows or columns of O, considered as row or column vectors, are orthonormal. All entries are bounded |Oij|≤1, and again this leads to O( n) being a compact group. The identity 1 = det (OTO) = detOTdetO= (detO)2(6.4) tells us that det O=±1. The subset of orthogonal matrices with det O= +1 constitute a subgroup of O( n) called the special orthogonal group , SO(n). The unimodularity condition discards a disconnected part of th e group manifold and does not reduce its dimension, which remains n(n−1)/2. 210 CHAPTER 6. LIE GROUPS 6.1.2 Symplectic Groups The symplectic groups (named from Greek meaning to “fold tog ether”) are probably less familiar than the other matrix groups. We start with a non-degenerate skew-symmetric matrix ω. The symplec- tic group Sp(2 n,F) is then defined by Sp(2n,F) ={S∈GL(2n,F) :STωS=ω}. (6.5) Here Fcan be RorC. When F=C, we still use the transpose “ T,” not†, in this definition. Setting S=I2n+/epsilon1Aand demanding that STωS=ωshows thatATω+ωA= 0. It does not matter what skew matrix ωwe start from, because we can always find a basis in which ωtakes its canonical form: ω=/parenleftbigg 0−In In0/parenrightbigg . (6.6) In this basis we find, after a short computation, that the most general form forAis A=/parenleftbigg a b c−aT/parenrightbigg . (6.7) Hereais anyn-by-nmatrix, and bandcare symmetric ( bT=band cT=c)n-by-nmatrices. If the matrices are real, then counting the degree s of freedom gives the dimension of the real symplectic group as dim Sp(2n,R) =n2+ 2×n 2(n+ 1) =n(2n+ 1). (6.8) The entries in a,b,c can be arbitrarily large. Sp(2 n,R) is not compact. The determinant of any symplectic matrix is +1. To see this ta ke the elements of ωto beωij, and let ω(x,y) =ωijxiyj(6.9) be the associated skew bilinear ( notsesquilinear) form . Then Weyl’s identity from exercise ??.??shows that Pf (ω) (detM) det|x1,...x 2n| =1 2nn!/summationdisplay π∈S2nsgn (π)ω(Mxπ(1),Mxπ(2))···ω(Mxπ(2n−1),Mxπ(2n)), 6.1. MATRIX GROUPS 211 for any linear map M. Ifω(x,y) =ω(Mx,My ), we conclude that det M= 1 — but preserving ωis exactly the condition that Mbe an element of the symplectic group. Since the matrices in Sp(2 n,F) are automatically unimodular there is no “special symplectic” group. Unitary Symplectic Group The intersection of two groups is also a group. We therefore d efine the unitary symplectic group as Sp(n) = Sp(2n,C)∩U(2n). (6.10) This group is compact. We will see that its dimension is n(2n+1), the same as the non-compact Sp(2 n,R). Sp(n) may also be defined as U( n,H) where Hdenotes the skew field of quaternions. Warning : Physics papers often make no distinction between Sp( n), which is a compact group, and Sp(2 n,R) which is non-compact. To add to the confusion the compact Sp( n) is also sometimes called Sp(2 n). You have to judge from the context what group the author has in mind. Physics Application: Kramers’ degeneracy. LetC=iˆσ2. Therefore C−1ˆσnC=−ˆσ∗ n. (6.11) A time-reversal invariant Hamiltonian containing L·Sspin-orbit interactions obeys C−1HC=H∗. (6.12) If we regard the 2 n-by-2nmatrixHas being an n-by-nmatrix whose entries Hijare themselves 2-by-2 matrices, which we expand as Hij=h0 ij+i3/summationdisplay n=1hn ijˆσn, then the condition (6.12) implies that the ha ijare real numbers. We say thatHisreal quaternionic . This is because the Pauli sigma matrices are algebraically isomorphic to Hamilton’s quaternions under the identification iˆσ1↔i, iˆσ2↔j, iˆσ3↔k.(6.13) 212 CHAPTER 6. LIE GROUPS The hermiticity of Hrequires that Hji=Hijwhere the overbar denotes quaternionic conjugation q0+iq1ˆσ1+iq2ˆσ2+iq3ˆσ3→q0−iq1ˆσ1−iq2ˆσ2−iq3ˆσ3. (6.14) IfHψ=Eψ, thenHCψ∗=Eψ∗. SinceCis skew,ψandCψ∗are necessarily orthogonal. Therefore all states are doubly degenerate. Th is isKramers’ degeneracy. Hmay be diagonalized by a matrix in U( n,H), where U( n,H) consists of those elements of U(2 n) that satisfy C−1UC=U∗. We may rewrite this condition as C−1UC=U∗⇒UCUT=C, so U(n,H) consists of the unitary matrices that preserve the skew mat rixC. Thus U(n,H)⊆Sp(n). Further investigation shows that U( n,H) = Sp(n). We can exploit the quaternionic viewpoint to count the dimen sions. Let U=I+/epsilon1Bbe in U(n,H), thenBij+Bji= 0. The diagonal elements of Bare thus pure “imaginary” quaternions having no part proportio nal toI. There are therefore 3 parameters for each diagonal element. The up per triangle has n(n−1)/2 independent elements, each with 4 parameters. Counting up , we find dim U(n,H) = dim Sp( n) = 3n+ 4×n 2(n−1) =n(2n+ 1). (6.15) Thus, as promised, we see that the compact group Sp( n) and the non- compact group Sp(2 n,R) have the same dimension. We can also count the dimension of Sp( n) by looking at our previous matrices A=/parenleftbigg a b c−aT/parenrightbigg whereabandcare now allowed to be complex, but with the restriction that S=I+/epsilon1Abe unitary. This requires Ato be skew-hermitian, so a=−a†, andc=−b†, whileb(and hence c) remains symmetric. There are n2free real parameters in a, andn(n+ 1) inb, so dim Sp(n) = (n2) +n(n+ 1) =n(2n+ 1) as before. 6.2. GEOMETRY OF SU(2) 213 Exercise 6.1 : Show that SO(2N)∩Sp(2N,R)∼=U(N). Hint: Group the 2 Nbasis vectors on which O(2 N) acts into pairs xnandyn, n= 1,...,N . Assemble these pairs into zn=xn+iynand¯z=xn−iyn. Let ωbe the linear map that takes xn→ynandyn→−xn. Show that the subset of SO(2N) that commutes with ωmixeszi’s only with zi’s and ¯zi’s only with ¯zi’s. 6.2 Geometry of SU(2) To get a sense of Lie groups as geometric objects, we will stud y the simplest non-trivial case of SU(2) in some detail. A general 2-by-2 complex matrix can be parametrized as U=/parenleftbigg x0+ix3ix1+x2 ix1−x2x0−ix3/parenrightbigg . (6.16) The determinant of this matrix is unity provided (x0)2+ (x1)2+ (x2)2+ (x3)2= 1. (6.17) When this condition is met, and if in addition the xiare real, the matrix is unitary:U†=U−1. The group manifold of SU(2) can therefore be identified with the three-sphere S3. We will take as local co-ordinates x1,x2,x3.When we desire to know x0we will find it from x0=/radicalbig 1−(x1)2−(x2)2−(x3)2. This co-ordinate chart only labels the points in the half of t he three-sphere withx0>0, but this is typical of any non-trivial manifold. A complet e atlas of charts can be constructed if needed. We can simplify our notation by using the Pauli sigma matrice s ˆσ1=/parenleftbigg 0 1 1 0/parenrightbigg ,ˆσ2=/parenleftbigg 0−i i0/parenrightbigg ,ˆσ3=/parenleftbigg 1 0 0−1/parenrightbigg . (6.18) These obey [ˆσi,ˆσj] = 2i/epsilon1ijkˆσk,andσi,ˆσj+ ˆσjˆσi= 2δijI. (6.19) In terms of them, we can write g=U=x0I+ix1ˆσ1+ix2ˆσ2+ix3ˆσ3. (6.20) 214 CHAPTER 6. LIE GROUPS Elements of the group in the neighbourhood of the identity di ffer frome≡I by real linear combinations of the iˆσi. The three-dimensional vector space spanned by these matrices is therefore the tangent space TGeat the identity element. For any Lie group this tangent space is called the Lie algebra , g= LieGof the group. There will be a similar set of matrices iˆλifor any matrix group. They are called the generators of the Lie algebra, and satisfy commutation relations of the form [iˆλi,iˆλj] =−fk ij(iˆλk), (6.21) or equivalently [ˆλi,ˆλj] =ifk ijˆλk (6.22) Thefk ijare called the structure constants of the algebra. The “ i”’s associ- ated with the ˆλ’s in this expression are conventional in physics texts beca use for quantum mechanics application we usually desire the ˆλito be hermitian. They are usually absent in books written for mathematicians . Exercise 6.2 : Let ˆλ1andˆλ2be hermitian matrices. Show that if we define ˆλ3 by the relation [ ˆλ1,ˆλ2] =iˆλ3, then ˆλ3is also a hermitian matrix. Exercise 6.3 : For the group O( n) the matrices “ iˆλ” are real n-by-nskew symmetric matrices A. Show that if A1andA2are real skew symmetric matrices, then so is [ A1,A2]. Exercise 6.4 : For the group Sp(2 n,R) theiˆλmatrices are of the form A=/parenleftbigga b c−aT/parenrightbigg whereais any real n-by-nmatrix and bandcare symmetric ( aT=aand bT=b) realn-by-nmatrices. Show that the commutator of any two matrices of this form is also of this form. 6.2.1 Invariant vector fields Consider a matrix group, and in it a group element I+i/epsilon1ˆλilying close to the identity e≡I. Draw an arrow connecting ItoI+i/epsilon1ˆλi, and regard this arrow as a vector Lilying inTGe. Next map the infinitesimal element I+i/epsilon1ˆλito the neighbourhood an arbitrary group element gby multiplying 6.2. GEOMETRY OF SU(2) 215 on the leftto getg(I+i/epsilon1ˆλi). By drawing an arrow from gtog(I+i/epsilon1ˆλi), we obtain a vector Li(g) lying inTGg. This vector at gis the push forward of the vector at eby left multiplication by g. For example, consider SU(2) with infinitesimal element I+i/epsilon1ˆσ3. We find g(I+i/epsilon1ˆσ3) = (x0+ix1ˆσ1+ix2ˆσ2+ix3ˆσ3)(I+i/epsilon1ˆσ3) = (x0−/epsilon1x3) +iˆσ1(x1−/epsilon1x2) +iˆσ2(x2+/epsilon1x1) +iˆσ3(x3+/epsilon1x0). (6.23) This computation can also be interpreted as showing that the multiplication ofg∈SU(2) on the rightby (I+i/epsilon1ˆσ3) displaces the point g, changing its xi parameters by an amount δ x0 x1 x2 x3 =/epsilon1 −x3 −x2 x1 x0 . (6.24) Knowing how the displacement looks in terms of the x1,x2,x3co-ordinate system lets us read off the ∂/∂xµcomponents of the vector L3lying inTGg L3=−x2∂1+x1∂2+x0∂3. (6.25) Sincegcan be any point in the group, we have constructed a globally d efined vector field L3that acts on a function F(g) on the group manifold as L3F(g) = lim /epsilon1→0/braceleftbigg1 /epsilon1[F(g(I+i/epsilon1ˆσ3))−F(g)]/bracerightbigg . (6.26) Similarly we obtain L1=x0∂1−x3∂2+x2∂3 L2=x3∂1+x0∂2−x1∂3. (6.27) The vector fields Liare said to be left invariant because the push-forward of the vector Li(g) lying in the tangent space at gby multiplication on the left by any g/primeproduces a vector g/prime ∗[Li(g)] lying in the tangent space at g/primeg, and this pushed-forward vector coincides with the Li(g/primeg) already there. We can express this statement tersely as g∗Li=Li. 216 CHAPTER 6. LIE GROUPS Using∂ix0=−xi/x0,i= 1,2,3, we can compute the Lie brackets and find [L1,L2] =−2L3. (6.28) In general [Li,Lj] =−2/epsilon1ijkLk, (6.29) which coincides with the matrix commutator of the iˆσi. This construction works for all Lie groups. For each basis ve ctorLiin the tangent space at the identity e, we push it forward to the tangent space at g by left multiplication by g, and so construct the global left-invariant vector fieldLi. The Lie bracket of these vector fields will be [Li,Lj] =−fk ijLk, (6.30) where the coefficients fk ijare guaranteed to be position independent because (see exercise 3.5) the operation of taking the Lie bracket of two vector fields commutes with the operation of pushing-forward the vector fi elds. Con- sequently the Lie bracket at any point is just the image of the Lie bracket calculated at the identity. When the group is a matrix group, this Lie bracket will coincide with the commutator of the iˆλi, that group’s analogue of the iˆσimatrices. The Exponential Map Recall that given a vector field X≡Xµ∂µwe define associated flowby solving the equation dxµ dt=Xµ(x(t)). (6.31) If we do this for the left-invariant vector field L, with initial condition x(0) =e, we obtain a t-dependent group element g(x(t)), which we denote by Exp (tL). The symbol “Exp ” stands for the exponential map which takes elements of the Lie algebra to elements of the Lie group. The r eason for the name and notation is that for matrix groups this operation co rresponds to the usual exponentiation of matrices. Elements of the matri x Lie group are therefore exponentials of matrices in the the Lie algebra. T o see this suppose thatLiis the left invariant vector field derived from iˆλi. Then the matrix g(t) = exp(itˆλi)≡I+itˆλi−1 2t2ˆλ2−i1 3!t3ˆλ3+··· (6.32) 6.2. GEOMETRY OF SU(2) 217 is an element of the group, and g(t+/epsilon1) = exp(itˆλ) exp(i/epsilon1ˆλi) =g(t)/parenleftBig I+i/epsilon1ˆλi+O(/epsilon12)/parenrightBig . (6.33) From this we deduce that d dtg(t) = lim /epsilon1→0/braceleftbigg1 /epsilon1[g(t)(I+i/epsilon1ˆλi)−g(t)]/bracerightbigg =Lig(t). (6.34) Since exp(itˆλ) =Iwhent= 0, we deduce that Exp ( tLi) = exp(itˆλi). Right-invariant vector fields We can use multiplication on the rightto push forward an infinitesimal group element. For example: (I+i/epsilon1ˆσ3)g= (I+i/epsilon1ˆσ3)(x0+ix1ˆσ1+ix2ˆσ2+ix3ˆσ3) = (x0−/epsilon1x3) +iˆσ1(x1+/epsilon1x2) +iˆσ2(x2−/epsilon1x1) +iˆσ3(x3+/epsilon1x0). (6.35) This motion corresponds to the right-invariant vector field R3=x2∂1−x1∂2+x0∂3. (6.36) Similarly, we obtain R1=x3∂1−x0∂2+x1∂3 R2=x0∂1+x3∂2−x2∂3, (6.37) and find that [R1,R2] = +2R3. (6.38) In general, [Ri,Rj] = +2/epsilon1ijkRk. (6.39) For any Lie group, the Lie brackets of the right-invariant fie lds will be [Ri,Rj] = +fijkRk. (6.40) whenever [Li,Lj] =−fijkLk, (6.41) 218 CHAPTER 6. LIE GROUPS are the Lie brackets of the left-invariant fields. The relati ve minus sign be- tween the bracket algebra of the left and right invariant vec tor fields has the same origin as the relative sign between the commutators of space- and body-fixed rotations in classical mechanics. Because multi plication from the left does not interfere with multiplication from the right, the left and right invariant fields commute: [Li,Rj] = 0. (6.42) 6.2.2 Maurer-Cartan Forms Ifg∈G, thendgg−1∈LieG. For example, starting from g=x0+ix1ˆσ1+ix2ˆσ2+ix3ˆσ3 g−1=x0−ix1ˆσ1−ix2ˆσ2−ix3ˆσ3 (6.43) we have dg=dx0+idx1ˆσ1+idx2ˆσ2+idx3ˆσ3 = (x0)−1(−x1dx1−x2dx2−x3dx3) +idx1ˆσ1+idx2ˆσ2+idx3ˆσ3. (6.44) From this we find dgg−1=iˆσ1/parenleftbig (x0+ (x1)2/x0)dx1+ (x3+ (x1x2)/x0)dx2+ (−x2+ (x1x3)/x0)dx3/parenrightbig +iˆσ2/parenleftbig (−x3+ (x2x1)/x0)dx1+ (x0+ (x2)2/x0)dx2+ (x1+ (x2x3)/x0)dx3/parenrightbig +iˆσ3/parenleftbig (x2+ (x3x1)/x0)dx1+ (−x1+ (x3x2)/x0)dx2+ (x0+ (x3)2/x0)dx3/parenrightbig . (6.45) The part proportional to the identity matrix has cancelled. The result is therefore a Lie algebra-valued 1-form. We define the (right i nvariant) Maurer- Cartan forms ωi Rby dgg−1=ωR= (iˆσi)ωi R. (6.46) If we evaluate one-form ω1 Ron the right invariant vector field R1, we find ω1 R(R1) = (x0+ (x1)2/x0)x0+ (x3+ (x1x2)/x0)x3+ (−x2+ (x1x3)/x0)(−x2) = (x0)2+ (x1)2+ (x2)2+ (x3)2 = 1. (6.47) 6.2. GEOMETRY OF SU(2) 219 Working similarly, we find ω1 R(R2) = (x0+ (x1)2/x0)(−x3) + (x3+ (x1x2)/x0)x0+ (−x2+ (x1x3)/x0)x1 = 0. (6.48) In general we discover that ωi R(Rj) =δi j. These Maurer-Cartan forms there- fore constitute the dual basis to the right-invariant vecto r fields. We may also define the left invariant Maurer-Cartan forms g−1dg=ωL= (iˆσi)ωi L. (6.49) These obey ωi L(Lj) =δi j, showing that the ωi Lare the dual basis to the left-invariant vector fields. Acting with the exterior derivative dongg−1=Itells us that d(g−1) = −g−1dgg−1. By exploiting this fact, together with the anti-derivatio n prop- erty d(a∧b) =da∧b+ (−1)pa∧db, we may compute the exterior derivative of ωR. We find that dωR=d(dgg−1) = (dgg−1)∧(dgg−1) =ωR∧ωR. (6.50) A matrix product is implicit here. If it were not, the product of the two identical 1-forms on the right would automatically be zero. If we make this matrix structure explicit we find that ωR∧ωR=ωi R∧ωj R(iˆσi)(iˆσj) =1 2ωi R∧ωj R[iˆσi,iˆσj] =−1 2fk ij(iˆσk)ωi R∧ωj R, (6.51) so dωk R=−1 2fk ijωi R∧ωj R. (6.52) These equations are known as the Maurer-Cartan relations for the right- invariant forms. For the left-invariant forms we have dωL=d(g−1dg) =−(g−1dg)∧(g−1dg) =−ωL∧ωL, (6.53) 220 CHAPTER 6. LIE GROUPS or dωk L= +1 2fk ijωi L∧ωj L. (6.54) The Maurer-Cartan relations appear when we quantize gauge t heories. They are one part of the BRST transformations of the Fadeev-P opov ghost fields. 6.2.3 Euler Angles In physics it is common to use Euler angles to parameterize SU(2). We can write an arbitrary SU(2) matrix Uas a product U= exp{−iφˆσ3/2}exp{−iθˆσ2/2}exp{−iψˆσ3/2}, =/parenleftbigg e−iφ/20 0eiφ/2/parenrightbigg/parenleftbigg cosθ/2−sinθ/2 sinθ/2 cosθ/2/parenrightbigg/parenleftbigg e−iψ/20 0eiψ/2/parenrightbigg , =/parenleftbigg e−i(φ+ψ)/2cosθ/2−ei(ψ−φ)/2sinθ/2 ei(φ−ψ)/2sinθ/2e+i(ψ+φ)/2cosθ/2/parenrightbigg . (6.55) Comparing with the earlier expression for Uin terms of the xµ, we obtain the Euler-angle parameterization of the three-sphere x0= cosθ/2 cos(ψ+φ)/2, x1= sinθ/2 sin(φ−ψ)/2, x2=−sinθ/2 cos(φ−ψ)/2, x3=−cosθ/2 sin(ψ+φ)/2. (6.56) If the angles are taken in the range 0 ≤φ<2π, 0≤θ <π , 0≤ψ <4πwe cover the entire three-sphere once. Exercise 6.5 : Show that the Hopf map, defined in chapter 3, Hopf : S3→S2 is the “forgetful” map ( θ,φ,ψ )→(θ,φ), whereθandφare spherical polar co-ordinates on the two-sphere. Exercise 6.6 : Show that U−1dU=−i 2ˆσiΩi L, where Ω1 L= sinψdθ−sinθcosψdφ, Ω2 L= cosψdθ−sinθsinψdφ, Ω3 L=dψ+ cosθdφ. 6.2. GEOMETRY OF SU(2) 221 Compare these 1-forms with the components ωX= sinψ˙θ−sinθcosψ˙φ, ωY= cosψ˙θ−sinθsinψ˙φ, ωZ=˙ψ+ cosθ˙φ. of the angular velocity ωof a body with respect to the body-fixedXYZ axes in the Euler-angle conventions of exercise 2.17. Similarly show that dUU−1=−i 2ˆσiΩi R, where Ω1 R=−sinφdθ+ sinθcosψdψ, Ω2 R= cosφdθ+ sinθsinψdψ, Ω3 R=dφ+ cosθdψ, Compare these 1-forms with components ωx,ωy,ωzof the same angular ve- locity vector ω, but now with respect to the space-fixed xyzframe. 6.2.4 Volume and Metric The manifold of any Lie group has a natural metric which is obt ained by transporting the Killing form (see section 6.3.2) from the t angent space at the identity to any other point gby either left or right multiplication by g. In the case of a compact group, the resultant left and right i nvariant metrics coincide. In the case of SU(2) this metric is the usua l metric on the three-sphere. Using the Euler angle expression for the xµto compute the dxµ, we can express the metric on the sphere as “ds2/prime/prime= (dx0)2+ (dx1)2+ (dx2)2+ (dx3)2, =1 4/parenleftbig dθ2+ cos2θ/2(dψ+dφ)2+ sin2θ/2(dψ−dφ)2/parenrightbig , =1 4/parenleftbig dθ2+dψ2+dφ2+ 2 cosθdφdψ/parenrightbig . (6.57) Here, to save space, we have used the traditional physics way of writing a metric. In the more formal notation, where we think of the met ric as being 222 CHAPTER 6. LIE GROUPS a bilinear function, we would write the last line as g(,) =1 4(dθ⊗dθ+dψ⊗dψ+dφ⊗dφ+ cosθ(dφ⊗dψ+dψ⊗dφ)) (6.58) From (6.58) we find g= det (gµν) =1 43/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0 1 cos θ 0 cosθ1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle =1 64(1−cos2θ) =1 64sin2θ. (6.59) The volume element,√gdθdφdψ , is therefore d(Volume) =1 8sinθdθdφdψ, (6.60) and the total volume of the sphere is Vol(S3) =1 8/integraldisplayπ 0sinθdθ/integraldisplay2π 0dφ/integraldisplay4π 0dψ= 2π2. (6.61) This coincides with the standard expression for the volume o fSd−1, the surface of the d-dimensional unit ball, Vol(Sd−1) =2πd/2 Γ(d 2), (6.62) whend= 4. Exercise 6.7 : Evaluate the Maurer-Cartan form ω3 Lin terms of the Euler angle parameterization and show that iω3 L=1 2tr (ˆσ3U−1dU) =−i 2(dψ+ cosθdφ). Now recall that the Hopf map takes the point on the three-sphe re with Euler angle co-ordinates ( θ,φ,ψ ) to the point on the two-sphere with spherical polar co-ordinates ( θ,φ). Thus, if we set A=−dψ−cosθdφ, then we find F≡dA= sinθdθdφ = Hopf∗(d[AreaS2]). 6.2. GEOMETRY OF SU(2) 223 Also observe that A∧F=−sinθdθdφdψ. From this show that Hopf index of the Hopf map itself is equal t o 1 16π2/integraldisplay S3A∧F=−1. Exercise 6.8 : Show that for Uthe defining two-by-two matrices of SU(2), we have /integraldisplay SU(2)tr [(U−1dU)3] = 24π2. Suppose we have a map g:R3→SU(2) such that g(x) goes to the identity element at infinity. Consider the integral S[g] =1 24π2/integraldisplay R3tr (g−1dg)3, where the 3-form tr ( g−1dg)3is the pull-back to R3of the form tr [( U−1dU)3] on SU(2). Show that if we vary g→g+δg, then δS[g] =1 24π2/integraldisplay R3d/braceleftBig 3tr/parenleftBig (g−1δg)(g−1dg)2/parenrightBig/bracerightBig = 0, and soS[g] is topological invariant of the map g. Conclude that the functional S[g] is an integer, that integer being the Brouwer degree, or win ding number, of the map g:S3→S3. Exercise 6.9 : Generalize the result of the previous problem to show, for a ny mappingx/mapsto→g(x) into a Lie group G, and fornan odd integer, that the n-form tr (g−1dg)nconstructed from the Maurer-Cartan form is closed, and that δtr (g−1dg)n=d/braceleftBig ntr/parenleftBig (g−1δg)(g−1dg)n−1/parenrightBig/bracerightBig . (Note that for even nthe trace of ( g−1dg)nvanishes identically.) 6.2.5 SO(3)/similarequalSU(2)/Z2 The groups SU(2) and SO(3) are locally isomorphic . They have the same Lie algebra, but differ in their global topology. Although ro tations in space are elements of SO(3), electrons respond to these rotations by transforming under the two-dimensional defining representation of SU(2) . As we shall see, 224 CHAPTER 6. LIE GROUPS this means that after a rotation through 2 πthe electron wavefunction comes back to minus itself. The resulting topological entangleme nt is characteristic of the spinor representation of rotations and is intimately connected wi th the Fermi statistics of the electron. The spin representati ons were discovered by´Elie Cartan in 1913, long before they were needed in physics. The simplest way to motivate the spin/rotation connection i s via the Pauli sigma matrices. These matrices are hermitian, tracel ess, and obey ˆσiˆσj+ ˆσjˆσi= 2δijI, (6.63) If, for anyU∈SU(2), we define ˆσ/prime i=UˆσiU−1, (6.64) then the ˆσ/prime iare also hermitian, traceless, and obey (6.63). Since the or iginal ˆσiform a basis for the space of hermitian traceless matrices, w e must have ˆσ/prime i= ˆσjRji (6.65) for some real 3-by-3 matrix having entries Rij. From (6.63) we find that 2δij= ˆσ/prime iˆσ/prime j+ ˆσ/prime jˆσ/prime i = (ˆσlRli)(ˆσmRmj) + (ˆσmRmj)(ˆσlRli) = (ˆσlˆσm+ ˆσmˆσl)RliRmj = 2δlmRliRmj. Thus RmiRmk=δik. (6.66) In other words, RTR=I, andRis an element of O(3). Now the determinant of any orthogonal matrix is ±1, but the manifold of SU(2) is a connected set andR=IwhenU=I. Since a continuous map from a connected set to the integers must be a constant, we conclude that det R= 1 for all U. The Rmatrices are therefore in SO(3). We now exploit the principle of the sextant to show that the co rrespon- dance goes both ways, i.e.we can find a U(R) for any element R∈SO(3). 6.2. GEOMETRY OF SU(2) 225 Left−hand half of fixed hand half is transparant mirror is silvered. Right− View through telescope of sun brought down to touch horizon 120 90 60300 o o ooo θPivotMovable Mirror 2θTo sun Telescope Fixed, half silvered mirrorTo Horizon Figure 6.1: The sextant. This familiar instrument is used to measure the altitude of t he sun above the horizon while standing on the pitching deck of a ship at sea. A theodolite or similar device would be rendered useless by the ship’s motio n. The sextant exploits the fact that successive reflection in two mirrors i nclined at an angle θto one another serves to rotate the image through an angle 2 θabout the line of intersection of the mirror planes. This rotation is u sed to superimpose the image of the sun onto the image of the horizon, where it sta ys even if the instrument is rocked back and forth. Exactly the same tri ck is used in constructing the spinor representations of the rotation gr oup. To do this, consider a vector xwith components xiand form the matrix /hatwidex=xiˆσi. Now, if nis a unit vector with components ni, then (−ˆσini)/hatwidex(ˆσknk) =/parenleftbig xj−2(n·x)(nj)/parenrightbig ˆσj=/hatwidex−2(n·x)/hatwiden (6.67) The vector x−2(n·x)nis the result of reflecting xin the plane perpendicular ton. Consequently −(ˆσ1cosθ/2 + ˆσ2sinθ/2)(−ˆσ1)/hatwidex(ˆσ1)(ˆσ1cosθ/2 + ˆσ2sinθ/2) (6.68) 226 CHAPTER 6. LIE GROUPS performs two successive reflections on x, first in the “1” plane, and then in a plane at an angle θ/2 to it. Multiplying out the factors, and using the ˆ σi algebra, we find (cosθ/2−ˆσ1ˆσ2sinθ/2)ˆx(cosθ/2 + ˆσ1ˆσ2sinθ/2) = ˆσ1(cosθx1−sinθx2) + ˆσ2(sinθx1+ cosθx2) + ˆσ3x3.(6.69) The effect on xis a rotation through θ, as claimed. We can drop the xiand re-express (6.69) as UˆσiU−1= ˆσjRji, (6.70) whereRijis the 3-by-3 rotation matrix for a rotation through angle θin the 1-2 plane, and U= exp/braceleftbigg −i 2ˆσ3θ/bracerightbigg = exp/braceleftbigg −i1 4i[ˆσ1,ˆσ2]θ/bracerightbigg (6.71) is an element of SU(2). We have exhibited two ways of writing t he exponents in (6.71) because the subscript 3 on ˆ σ3indicates the axis about which we are rotating, while the 1 ,2 in [ˆσ1,ˆσ2] indicates the plane in which the rotation occurs. It is the second language that generalizes to higher dimensions. More on the use of mirrors for creating and combining rotations ca n be found in the the appendix to Misner, Thorn, and Wheeler’s Gravitation . The mirror construction shows that for any R∈SO(3) there is a two- dimensional unitary matrix U(R) such that U(R)ˆσiU−1(R) = ˆσjRji. (6.72) ThisU(R) is not unique however. If U∈SU(2) then so is−U. Furthermore U(R)ˆσiU−1(R) = (−U(R))ˆσi(−U(R))−1, (6.73) and soU(R) and−U(R) implement exactly the same rotation R. Conversely, if two SU(2) matrices U,Vobey UσiU−1=VσiV−1(6.74) thenV−1Ucommutes with all 2-by-2 matrices and, by Schur’s lemma, mus t be a multiple of the identity. But if λI∈SU(2) then λ=±1. ThusU=±V. The mapping between SU(2) and SO(3) is therefore two-to-one . SinceUand 6.2. GEOMETRY OF SU(2) 227 −Ucorrespond to the same R, the group manifold of SO(3) is the three- sphere with antipodal points identified . Unlike the two-sphere, where the identification of antipodal points gives the non-orientabl e projective plane, this three-manifold is is orientable. It is not, however, si mply connected: a path on the three-sphere from a point to its antipode forms a c losed loop in SO(3), but one not contractable to a point. If we continue o n from the antipode back to the original point, the combined path iscontractable. This means that the first Homotopy group , the group of based paths with composi- tion given by concatenation, is π1(SO(3)) = Z2. This is the topology behind the Phillipine (or Balinese) Candle Dance, and is how the ele ctron knows whether a sequence of rotations that eventually bring it bac k to its original orientation should be counted as a 360◦rotation (U=−I) or a 720◦∼0◦ rotation (U= +I). Exercise 6.10 : Verify that U(R)ˆσiU−1(R) = ˆσjRji is consistent with U(R2)U(R1) =±U(R2R1). Spinor representations of SO(N) The mirror trick can be extended to perform rotations in Ndimensions. We replace the three ˆ σimatrices by a set of NDirac gamma matrices , which obey the defining relations of a Clifford algebra ˆγµˆγν+ ˆγνˆγµ= 2δµνI. (6.75) These relations are a generalization of the key algebraic pr operty of the Pauli sigma matrices. IfN(= 2n) is even, then we can find 2n-by-2nhermitian matrices, ˆ γµ, satisfying this algebra. If N(= 2n+1) is odd, we append to the matrices for N= 2nthe hermitian matrix ˆ γ2n+1=−(i)nˆγ1ˆγ2···ˆγ2nwhich obeys ˆ γ2 2n+1= 1 and anti-commutes with all the other ˆ γµ. The ˆγmatrices therefore act on a 2[N/2]dimensional space, where the square brackets denote the integer part ofN/2. The ˆγ’s do not form a Lie algebra as they stand, but a rotation throu gh θin themn-plane is obtained from e−i1 4i[ˆγm,ˆγn]θˆγiei1 4i[ˆγm,ˆγn]θ= ˆγjRji, (6.76) 228 CHAPTER 6. LIE GROUPS and we find that the hermitian matrices ˆΓmn=1 4i[ˆγm,ˆγn] form a basis for the Lie algebra of SO( N). The 2[N/2]dimensional space on which they act is the Dirac spinor representation of SO( N). Although the matrices exp {iˆΓµνθµν} are unitary, they are not, in general, the entirety of U(2[N/2]), but instead constitute a subgroup called Spin( N). IfNis even then we can still construct the matrix ˆ γ2n+1that anti- commutes with all the other ˆ γµ’s. It cannot be the identity matrix, therefore, but it commutes with all the Γ mn. By Schur’s lemma, this means that the SO(2n) Dirac spinor representation space Visreducible . Now ˆγ2 2n+1=I, and so ˆγ2n+1has eigenvalues±1. The two eigenspaces are invariant under the action of the group, and thus the Dirac spinor space decom poses into two irreducible Weyl spinor representations V=Vodd⊕Veven. (6.77) HereVevenandVodd, the plus and minus eigenspaces of ˆ γ2n+1, are called the spaces of right and left chirality . WhenNis odd the spinor representation is irreducible. Exercise 6.11 : Starting from the defining relations of the Clifford algebra (6.75) show that, for N= 2n, tr (ˆγµ) = 0, tr (ˆγ2n+1) = 0, tr (ˆγµˆγν) = tr (I)δµν, tr (ˆγµˆγνˆγσ) = 0, tr (ˆγµˆγνˆγσˆγτ) = tr (I)(δµνδστ−δµσδντ+δµτδνσ). Exercise 6.12 : Consider the space Ω( C) =/circleplustext pΩp(C) of complex-valued skew symmetric tensors Aµ1...µpfor 0≤p≤N= 2n. Let ψαβ=N/summationdisplay p=01 p!/parenleftbig ˆγµ1···ˆγµp/parenrightbig αβAµ1...µp define a mapping from Ω( C) into the space of complex matrices of the same size as the ˆγµ. Show that this mapping is invertible — i.e.givenψαβyou can recover the Aµ1...µp. By showing that the dimension of Ω( C) is 2N, deduce that the ˆγµmust be at least 2n-by-2nmatrices. 6.2. GEOMETRY OF SU(2) 229 Exercise 6.13 : Show that the R2nDirac operator D= ˆγµ∂µobeysD2=∇2. Recall that Hodge operator d−δfrom section 4.7.1 is also a “square root” of the Laplacian: (d−δ)2=−(dδ+δd) =∇2. Show that ψαβ→(Dψ)αβ= (ˆγµ)αα/prime∂µψα/primeβ corresponds to the action of d−δon the space Ω( R2n,C) of differential forms A=1 p!Aµ1...µp(x)dxµ1···dxµp The space of complex-valued differential forms has thus been made to look like a collection of 2nDirac spinor fields, one for each value of the “flavour index” β. Theseψαβare called K¨ ahler-Dirac fields. They are not really flavoured spinors because a rotation transforms both the αandβindices. Exercise 6.14 : That a set of 2 nDiracγ’s have a 2n-by-2nmatrix representation is most naturally established by using the tools of second qu antization. To this end, letai,a† ii= 1,...,n be set of anti-commuting annihilation and creation operators obeying aiaj+ajai= 0, aia† j+a† jai=δijI, and let|0/angbracketrightbe the “no particle” state such that ai|0/angbracketright= 0,i= 1,...,n . Then the 2nstates |m1,...,mn/angbracketright= (a† 1)m1···(a† n)mn|0/angbracketright, where themitake the value 0 or 1, constitute a basis for a space on which th e aianda† iact irreducibly. Show that the 2 noperators γi=ai+a† i γi+n=i(ai−a† i) obey γµγν+γνγµ= 2δµνI, and hence can be represented by 2n-by-2nmatrices. Deduce further that spaces specs of left and right chirality are the spaces of odd or even “particle number.” 230 CHAPTER 6. LIE GROUPS The Adjoint Representation The spin/rotation correspondence involves conjugation: ˆ σi→UˆσiU−1. The idea of obtaining a representation by conjugation works for an arbitrary Lie group. It is easiest, however, to describe in the case of a mat rix group where we consider an infinitesimal element I+i/epsilon1ˆλi. The conjugate element g(I+i/epsilon1ˆλi)g−1will also be an infinitesimal element. Since gIg−1=I, this means that g(iˆλi)g−1must be expressible as a linear combination of the iˆλi matrices. Consequently we can define a linear map acting on th e element X=ξiˆλiof the Lie algebra by setting Ad(g)ˆλi≡gˆλig−1=ˆλj[Ad (g)]j i. (6.78) The matrices with entries [Ad( g)]j iform the adjoint representation of the group. The dimension of the adjoint representation coincid es with that of the group manifold. The spinor construction shows that the d efining repre- sentation of SO(3) is the adjoint representation of SU(2). For a general Lie group, we make Ad( g) act on a vector in the tangent space at the identity by pushing the vector forward to TGgby left multiplica- tion byg, and then pushing it back from TGgtoTGeby right multiplication byg−1. Exercise 6.15 : Show that [Ad (g1g2)]j i= [Ad (g1)]j k[Ad (g2)]k i, thus confirming that Ad( g) is a representation. 6.2.6 Peter-Weyl Theorem The volume element constructed in section 6.2.4 has the feat ure that it is invariant. In other words if we have a subset Ω of the group manifold with volumeV, then the image set gΩ under left multiplication has the exactly the same volume. We can also construct a volume element that is in variant under right multiplication by g, and in general these will be different. For a group whose manifold is a compact set, however, both left- and righ t-invariant volume elements coincide. The resulting measure on the grou p manifold is called the Haar measure. For acompact group, therefore, we can replace the sums over the group elements that occur in the representation theory of finite gr oups, by con- vergent integrals over the group elements using the invaria nt Haar measure, 6.2. GEOMETRY OF SU(2) 231 which is usually denoted by d[g] . The invariance property is expressed by d[g1g] =d[g] for any constant element g1. This allows us to make a change- of-variables transformation, g→g1g, identical to that which played such an important role in deriving the finite group theorems. Conseq uently, all the results from finite groups, such as the existence of an invari ant inner product and the orthogonality theorems, can be taken over by the simp le replacement of a sum by an integral. In particular, if we normalize the mea sure so that the volume of the group manifold is unity, we have the orthogo nality relation /integraldisplay d[g]/parenleftbig DJ ij(g)/parenrightbig∗DK lm(g) =1 dimJδJKδilδjm. (6.79) The Peter-Weyl theorem asserts that the representation mat rices,DJ mn(g), form a complete set of orthogonal functions on the group mani fold. In the case of SU(2) this tells us that the spin Jrepresentation matrices DJ mn(θ,φ,ψ ) =/angbracketleftJ,m|e−iJ3φe−iJ2θe−iJ3ψ|J,n/angbracketright, =e−imφdJ mn(θ)e−inψ, (6.80) which you will know from quantum mechanics courses,1are a complete set of functions on the three-sphere with 1 16π2/integraldisplayπ 0sinθdθ/integraldisplay2π 0dφ/integraldisplay4π 0dψ/parenleftbig DJ mn(θ,φ,ψ )/parenrightbig∗DJ/prime m/primen/prime(θ,φ,ψ ) =1 2J+ 1δJJ/primeδmm/primeδnn/prime. (6.81) Since theDL m0(whereLhas to be an integer for n= 0 to be possible) are independent of the third Euler angle, ψ, we can do the trivial integral over ψto get 1 4π/integraldisplayπ 0sinθdθ/integraldisplay2π 0dφ/parenleftbig DL m0(θ,φ)/parenrightbig∗DL/prime m/prime0(θ,φ) =1 2L+ 1δLL/primeδmm/prime.(6.82) Comparing with the definition of the spherical harmonics, we see that we can identify YL m(θ,φ) =/radicalbigg 2L+ 1 4π/parenleftbig DL m0(θ,φ,ψ )/parenrightbig∗. (6.83) 1See, for example, G. Baym Lectures on Quantum Mechanics , Ch 17. 232 CHAPTER 6. LIE GROUPS The complex conjugation is necessary here because DJ mn(θ,φ,ψ )∝e−imφ, whileYL m(θ,φ)∝eimφ. The character, χJ(g) =/summationtext nDJ nn(g) will be a function only of the angle θ we have rotated through, not the axis of rotation — all rotati ons through a common angle being conjugate to one another. Because of this χJ(θ) can be found most simply by looking at rotations about the zaxis, since these give rise to easily computed diagonal matrices. We find χ(θ) =eiJθ+ei(J−1)θ+···+e−i(J−1)θ+e−iJθ, =sin(2J+ 1)θ/2 sinθ/2. (6.84) Warning : The angle θin this formula and the next is the not the Euler angle. For integer J, corresponding to non-spinor rotations, a rotation throug h an angleθabout an axis nand a rotation though an angle 2 π−θabout−n are the same operation. The maximum rotation angle is theref oreπ. For spinor rotations this equivalence does not hold, and the rot ation angle θruns from 0 to 2 π. The character orthogonality must therefore be 1 π/integraldisplay2π 0χJ(θ)χJ/prime(θ) sin2/parenleftbiggθ 2/parenrightbigg dθ=δJJ/prime, (6.85) implying that the volume fraction of the rotation group cont aining rotations through angles between θandθ+dθis sin2(θ/2)dθ/π. Exercise 6.16 : Prove this last statement about the volume of the equivalen ce classes by showing that the volume of the unit three-sphere t hat lies between a rotation angle of θandθ+dθis 2πsin2(θ/2)dθ. 6.2.7 Lie Brackets vs.Commutators There is an irritating minus sign problem that needs to be ack nowledged. The Lie bracket [ X,Y] of of two vector fields is defined by first running along X, thenYand then back in the reverse order. If we do this for the action of matrices, ˆXandˆY, on a vector space, however, then, reading from right to left as we always do for matrix operations, we have e−t2ˆYe−t1ˆXet2ˆYet1ˆX=I−t1t2[ˆX,ˆY] +···, (6.86) 6.2. GEOMETRY OF SU(2) 233 which has the other sign. Consider for example rotations abo ut thex,y,z axes, and look at effect these have on the co-ordinates of a poi nt: Lx:/braceleftbigg δy=−zδθx δz= +yδθx/bracerightbigg =⇒Lx=y∂z−z∂y,ˆLx= 0 0 0 0 0−1 0 1 0 , Ly:/braceleftbigg δz=−xδθy δx= +zδθy/bracerightbigg =⇒Ly=z∂x−x∂z,ˆLy= 0 0 1 0 0 0 −1 0 0 , Lz:/braceleftbigg δx=−yδθz δy= +xδθz/bracerightbigg =⇒Lz=x∂y−y∂x,ˆLy= 0−1 0 1 0 0 0 0 0 . From this we find [Lx,Ly] =−Lz, (6.87) as a Lie bracket of vector fields, but [ˆLx,ˆLy] = + ˆLz, (6.88) as a commutator of matrices. This is the reason why it is the leftinvariant vector fields whose Lie bracket coincides with the commutato r of theiˆλi matrices. Some insight into all this can be had by considering the actio n of the left invariant fields on the representation matrices, DJ mn(g). For example LiDJ mn(g) = lim /epsilon1→0/bracketleftbigg1 /epsilon1/parenleftBig DJ mn(g(1 +i/epsilon1ˆλi))−DJ mn(g)/parenrightBig/bracketrightbigg = lim /epsilon1→0/bracketleftbigg1 /epsilon1/parenleftBig DJ mn/prime(g)DJ n/primen(1 +i/epsilon1ˆλi)−DJ mn(g)/parenrightBig/bracketrightbigg = lim /epsilon1→0/bracketleftbigg1 /epsilon1/parenleftBig DJ mn/prime(g)(δn/primen+i/epsilon1(ˆΛJ i)n/primen)−DJ mn(g)/parenrightBig/bracketrightbigg =DJ mn/prime(g)(iˆΛJ i)n/primen (6.89) where ˆΛJ iis the matrix representing ˆλiin the representation J. Repeating this exercise we find that Li/parenleftbig LjDJ mn(g)/parenrightbig =DJ mn/prime/prime(g)(iˆΛJ i)n/prime/primen/prime(iˆΛJ j)n/primen, (6.90) 234 CHAPTER 6. LIE GROUPS Thus [Li,Lj]DJ mn(g) =DJ mn/prime(g)[iˆΛJ i,iˆΛJ j]n/primen, (6.91) and we get the commutator of the representation matrices in t he “correct” order only if we multiply the infinitesimal elements in succe ssively from the right. There appears to be no escape from this sign problem. Many tex ts simply ignore it, a few define the Lie bracket of vector fields with the opposite sign, and a few simply point out the inconvenience and get on the wit h the job. We will follow the last route. 6.3 Lie Algebras A Lie algebra gis a (real or complex) finite-dimensional vector space with a non-associative binary operation g×g→gthat assigns to each ordered pair of elements, X1,X2, a third element called the Lie bracket, [ X1,X2]. The bracket is: a) Skew symmetric: [ X,Y] =−[Y,X], b) Linear: [ λX+µY,Z] =λ[X,Z] +µ[Y,Z], and in place of associativity, obeys c) The Jacobi identity: [[ X,Y],Z] + [[Y,Z],X] + [[Z,X],Y] = 0. Example: LetM(n) denote the algebra of real n-by-nmatrices. As a vector space over R, this algebra is n2dimensional. Setting [ A,B] =AB−BA, makesM(n) into a Lie Algebra. Example: Letb+denote the subset of M(n) consisting of upper triangular matrices with any number (including zero) allowed on the dia gonal. Then b+with the above bracket is a Lie algebra. (The “b” stands for th e French mathematician and statesman ´Emile Borel). Example: Letn+denote the subset of b+consisting of strictly upper trian- gular matrices — those with zero on the diagonal. Then n+with the above bracket is a Lie algebra. (The “n” stands for nilpotent. ) Example: LetGbe a Lie group, and Lithe left invariant vector fields. We know that [Li,Lj] =fk ijLk (6.92) where [,] is the Lie bracket of vector fields. The resulting Lie algebr a, g= LieGis the Lie algebra of the group. 6.3. LIE ALGEBRAS 235 Example: The setN+of upper triangular matrices with 1’s on the diagonal forms a Lie group and has n+as its Lie algebra. Similarly, the set B+ consisting of upper triangular matrices, with any non-zero number allowed on the diagonal, is also a Lie group, and has b+as its Lie algebra. Ideals and Quotient algebras As we saw in the examples, we can define subalgebras of a Lie alg ebra. If we want to define quotient algebras by analogy to quotient gro ups, we need a concept analogous to that of invariant subgroups. This is p rovided by the notion of an ideal. A ideal is a subalgebra i⊆gwith the property that [i,g]⊆i. (6.93) In other words, taking the bracket of any element of gwith any element ofigives an element in i. With this definition we can form g−iby identifying X∼X+Ifor anyI∈i. Then [X+i,Y+i] = [X,Y] +i, (6.94) and the bracket of two equivalence classes is insensitive to the choice of representatives. If a Lie group Ghas an invariant subgroup Hwhich is also a Lie group, then the Lie algebra hof the subgroup is an ideal in g= LieGand the Lie algebra of the quotient group G/H is the quotient algebra g−h. If the Lie algebra has no non-trivial ideals, then it is said t o besimple . The Lie algebra of a simple Lie group will be simple. Exercise 6.17 : Let i1andi2be ideals in g. Show that i1∩i2is also an ideal in g. 6.3.1 Adjoint Representation Given an element X∈glet it act on the Lie algebra considered as a vector space by a linear map ad ( x) defined by ad(X)Y= [X,Y]. (6.95) The Jacobi identity is then equivalent to the statement (ad (X)ad(Y)−ad(Y)ad(X))Z= ad ([X,Y])Z. (6.96) 236 CHAPTER 6. LIE GROUPS Thus (ad (X)ad(Y)−ad(Y)ad(X)) = ad ([X,Y]), (6.97) or [ad (X),ad(Y)] = ad ([X,Y]), (6.98) and the map X→ad (X) is a representation of the algebra called the adjoint representation . The linear map “ad ( X)” exponentiates to give a map exp[ad ( tX)] defined by exp[ad(tX)]Y=Y+t[X,Y] +1 2t2[X,[X,Y]] +···. (6.99) You probably know the matrix identity2 etABe−tA=B+t[A,B] +1 2t2[A,[A,B]] +···. (6.100) Now, earlier in the chapter, we defined the adjoint represent ation “Ad ” of thegroup on the vector space of the Lie algebra. We did this setting gXg−1= Ad (g)X. Comparing the two previous equations we see that Ad (ExpY) = exp(ad( Y)). (6.101) 6.3.2 The Killing form Using “ad” we can define an inner product /angbracketleft,/angbracketrighton a real Lie algebra by setting /angbracketleftX,Y/angbracketright= tr(ad (X)ad(Y)). (6.102) This inner product is called the Killing form , after Wilhelm Killing. Using the Jacobi identity, and the cyclic property of the trace, we find that /angbracketleftad(X)Y,Z/angbracketright+/angbracketleftY,ad(X)Z/angbracketright= 0, (6.103) or, equivalently, /angbracketleft[X,Y],Z/angbracketright+/angbracketleftY,[X,Z]/angbracketright= 0. (6.104) From this we deduce (by differentiating with respect to t) that /angbracketleftexp(ad(tX))Y,exp(ad(tX))Z/angbracketright=/angbracketleftY,Z/angbracketright, (6.105) 2In case you do not, it is easily proved by setting F(t) =etABe−tA, noting that d dtF(t) = [A,F(t)], and observing that the RHS is the unique series solution t o this equation satisfying the boundary condition F(0) =B. 6.3. LIE ALGEBRAS 237 so the Killing form is invariant under the action of the adjoi nt representation of the group on the algebra. When our group is simple, any other invariant inner product will be proportional to this Killing-form pro duct. Exercise 6.18 : Let ibe an ideal in g. Show that for I1,I2∈i /angbracketleftI1,I2/angbracketrightg=/angbracketleftI1,I2/angbracketrighti where/angbracketleft,/angbracketrightiis the Killing form on iconsidered as a Lie algebra in its own right. (This equality of inner products is not true for subal gebras that are not ideals.) Semi-simplicity Recall that a Lie algebra containing no non-trivial ideals i s said to be simple . When the Killing form is non degenerate, the Lie Algebra is sa id to be semi- simple . The reason for this name is that a semi-simple algebra is almost simple, in that it can be decomposed into a direct sum of decou pled simple algebras g=s1⊕s2⊕···⊕ sn. (6.106) Here the direct sum symbol “ ⊕” implies not only a direct sum of vector spaces but also that [ si,sj] = 0 fori/negationslash=j. The Lie algebra of all the matrix groups O( n), Sp(n), SU(n),etc.are semi-simple (indeed they are usually simple) but this is not true of the alge- brasn+andb+. Cartan showed that our Killing-form definition of semi-simp licity is equiv- alent his original definition of a Lie algebra being semi-sim ple if it contains noabelian ideal — i.e.no ideal with [ Ii,Ij] = 0 for all Ii∈i. The following exercises establish the direct sum decomposition, and, en passant , the easy half of Cartan’s result. Exercise 6.19 : Use the identity (6.104) to show that if i⊂gis an ideal, then i⊥, the set of elements orthogonal to iwith respect to the Killing form, is also an ideal. Exercise 6.20 : Show that if ais an abelian ideal, then every element of ais Killing perpendicular to the entire Lie algebra. (Thus non- degeneracy⇒no non-trivial abelian ideal. The null space of the Killing for m is not necessarily an abelian ideal, though, so establishing the converse is ha rder.) 238 CHAPTER 6. LIE GROUPS Exercise 6.21 : Letgbe semi-simple and i⊂gan ideal. We know from exercise 6.17 that i∩i⊥is an ideal. Use (6.104) coupled with the non-degeneracy of the Killing form to show that it is an abelian ideal. Use the previous exercise to conclude that i∩i⊥={0}, and from this that [ i,i⊥] = 0. Exercise 6.22 : Let/angbracketleft,/angbracketrightbe a non-degenerate inner product on a vector space V. LetW⊆Vbe a subspace. Show that dimW+ dimW⊥= dimV. (This is not as obvious as it looks. For a non-positive-defini te inner product WandW⊥can have a non-trivial intersection. Consider two-dimensi onal Minkowski space. If Wis the space of right-going, light-like, vectors then W≡W⊥, but dimW+ dimW⊥still equals two.) Exercise 6.23 : Put the two preceding exercises together to show that g=i⊕i⊥. Show that iandi⊥are semi-simple in their own right as Lie algebras. We can therefore continue to break up iandi⊥until we end with gdecomposed into a direct sum of simple algebras. Compactness If the Killing form is negative definite, a real Lie Algebra is said to be com- pact, and is the Lie algebra of a compact group. With the physicist ’s habit of writingiXifor the generators of the Lie algebra, a compact group has Killing metric tensor gij= tr{ad(Xi)ad (Xj)} (6.107) that is a positive definite matrix. In a basis where gij=δij, the exp(ad X) matrices of the adjoint representations of a compact group Gform a subgroup of the orthogonal group O( N), whereNis the dimension of G. Totally anti-symmetric structure constants Given a basis iXifor the Lie-algebra vector space, we define the structure constantsfijkby [Xi,Xj] =ifijkXk. (6.108) 6.3. LIE ALGEBRAS 239 In terms of the fijk, the skew symmetry of ad ( Xi), as expressed by equation (6.103), becomes 0 =/angbracketleftad (Xk)Xi,Xj/angbracketright+/angbracketleftXi,ad(Xk)Xj/angbracketright ≡ /angbracketleft[Xk,Xi],Xj/angbracketright+/angbracketleftXi,[Xk,Xj]/angbracketright =i(fkilglj+gilfkjl) =i(fkij+fkji). (6.109) In the last line we have used the Killing metric to “lower” the indexland so define the symbol fijk. Thusfijkis skew symmetric under the interchange of its second pair of indices. Since the skew symmetry of the L ie bracket ensures that fijkis skew symmetric under the interchange of the first pair of indices, it follows that fijkis skew symmetric under the interchange of any pair of its indices. By comparing the definition of the structure constants with [Xi,Xj] = ad (Xi)Xj=Xk[ad (Xi)]k j, (6.110) we read-off that the matrix representing ad( Xi) has entries [(ad(Xi)]k j=ifijk. (6.111) Consequently gij= tr{ad (Xi)ad(Xj)}=−fiklfjlk. (6.112) The quadratic Casimir The only “product” that is defined in the abstract Lie algebra gis the Lie bracket [X,Y]. Once we have found matrices forming a representation of the Lie algebra, however, we can form the ordinary matrix pro duct of these. Suppose that we have a Lie algebra gwith basisXiand have found matrices ˆXiwith the same commutation relations as the Xi. Suppose further that the algebra is semisimple and so gij, the inverse of the Killing metric, exists. We can usegijto construct the matrix ˆC2=gijˆXiˆXj. (6.113) This matrix is called the quadratic Casimir operator, after Hendrik Casimir. Its chief property is that it commutes with all the ˆXi: [ˆC2,ˆXi] = 0. (6.114) 240 CHAPTER 6. LIE GROUPS If our representation is irreducible then Shur’s lemma tell s us that ˆC2=c2I (6.115) where the number c2is referred to as the “value” of the quadratic Casimir in that irrep.3 Exercise 6.24 : Show that [ ˆC2,Xi] = 0 is another consequence of the complete skew symmetry of the fijk. 6.3.3 Roots and Weights We now want to study the representation theory of Lie groups. It is, in fact, easier to study the representations of the Lie algebra, and t hen exponentiate these to find the representations of the group. In other words given an abstract Lie algebra with bracket [Xi,Xj] =ifijkXk, (6.116) we seek to find all matrices ˆXJ isuch that [ˆXJ i,ˆXJ j] =ifijkˆXJ k. (6.117) (Here, as with the representations of finite groups, we use th e superscript Jto distinguish one representation from another.) Then, given a representation ˆXJ iof the Lie algebra, the matrices DJ(g(ξ)) = exp/braceleftBig iξiˆXJ i/bracerightBig , (6.118) whereg(ξ) = Exp{iξiXi}, will form a representation of the Lie group. To be more precise, they will form a representation of that part of the group which is connected to the identity element. The numbers ξiwill serve as co-ordinates for some neighbourhood of the identity. For co mpact groups there will be a restriction on the range of the ξibecause there must be ξifor which exp/braceleftBig iξiˆXJ i/bracerightBig =I. 3Mathematicians do sometimes consider formal products of Li e algebra elements X,Y∈ g. When they do, they equip them with the rule that XY−YX−[X,Y] = 0,whereXY andYXare formal products, and [ X,Y] is the Lie algebra product. These formal products are not elements of the Lie algebra, but instead live in an ext ended mathematical structure called the Universal enveloping algebra ofg, and denoted by U(g). The quadratic Casimir can then be considered to be an element of this larger algebra . 6.3. LIE ALGEBRAS 241 SU(2) The quantum-mechanical angular momentum algebra consists of the com- mutation relation [J1,J2] =i/planckover2pi1J3, (6.119) together with two similar equations related by cyclic permu tations. This, once we set /planckover2pi1= 1, is the Lie algebra su(2) of the group SU(2). The goal of representation theory is to find all possible sets of matri ces which have the same commutation relations as these operators. Since th e group SU(2) is compact, we can use the group-averaging trick from section 5 .2.2 to define an inner product with respect to which these representations a re unitary, and the matrices Jihermitian. Remember how this problem is solved in quantum mechanics cou rses, where we find a representation for each spin j=1 2,1,3 2,etc.We begin by constructing “ladder” operators J+=J1+iJ2, J −=J† +=J1−iJ2, (6.120) which are eigenvectors of ad ( J3) ad (J3)J±= [J3,J±] =±J±. (6.121) From (6.121) we see that if |j,m/angbracketrightis an eigenstate of J3with eigenvalue m, thenJ±|j,m/angbracketrightis an eigenstate of J3with eigenvalue m±1. Now in any finite-dimensional representation there must be a highest weight state,|j,j/angbracketright, such that J3|j,j/angbracketright=j|j,j/angbracketrightfor some real number j, and such thatJ+|j,j/angbracketright= 0. From|j,j/angbracketrightwe work down by successive applications ofJ−to find|j,j−1/angbracketright,|j,j−2/angbracketright...We can find the normalization factors of the states|j,m/angbracketright∝(J−)j−m|j,j/angbracketrightby repeated use of the identities J+J−= (J2 1+J2 2+J2 3)−(J2 3−J3), J−J+= (J2 1+J2 2+J2 3)−(J2 3+J3). (6.122) The combination J2≡J2 1+J2 2+J2 3is the quadratic Casimir of su(2), and hence in any irrep is proportional to the identity matrix: J2=c2I. Because 0 =/bardblJ+|j,j/angbracketright/bardbl2 =/angbracketleftj,j|J† +J+|j,j/angbracketright =/angbracketleftj,j|J−J+|j,j/angbracketright =/angbracketleftj,j|/parenleftbig J2−J3(J3+ 1)/parenrightbig |j,j/angbracketright = [c2−j(j+ 1)]/angbracketleftj,j|j,j/angbracketright, (6.123) 242 CHAPTER 6. LIE GROUPS and/angbracketleftj,j|j,j/angbracketright≡/bardbl|j,j/angbracketright/bardbl2is not zero, we must have c2=j(j+ 1). We now compute /bardblJ−|j,m/angbracketright/bardbl2=/angbracketleftj,m|J† −J−|j,m/angbracketright =/angbracketleftj,m|J+J−|j,m/angbracketright =/angbracketleftj,m|/parenleftbig J2−J3(J3−1)/parenrightbig |j,m/angbracketright = [j(j+ 1)−m(m−1)]/angbracketleftj,m|j,m/angbracketright, (6.124) and deduce that the resulting set of normalized states |j,m/angbracketrightcan be chosen to obey J3|j,m/angbracketright=m|j,m/angbracketright, J−|j,m/angbracketright=/radicalbig j(j+ 1)−m(m−1)|j,m−1/angbracketright, J+|j,m/angbracketright=/radicalbig j(j+ 1)−m(m+ 1)|j,m+ 1/angbracketright. (6.125) If we takejto be an integer or a half-integer, we will find that J−|j,−j/angbracketright= 0. In this case we are able to construct a total of 2 j+ 1 states, one for each integer-spaced min the range−j≤m≤j. If we select some other fractional value forj, then the set of states will not terminate gracefully, and we will find an infinity of states with m<−j. These will have /bardblJ−|j,m/angbracketright/bardbl2<0, so the resultant representation cannot be unitary. SU(3) The strategy of finding ladder operators works for any semi-s imple Lie al- gebra. Consider, for example, su(3) = Lie(SU(3)). The matrix Lie algebra su(3) is spanned by the Gell-Mann λ-matrices ˆλ1= 0 1 0 1 0 0 0 0 0 ,ˆλ2= 0−i0 i0 0 0 0 0 ,ˆλ3= 1 0 0 0−1 0 0 0 0 , ˆλ4= 0 0 1 0 0 0 1 0 0 ,ˆλ5= 0 0−i 0 0 0 i0 0 ,ˆλ6= 0 0 0 0 0 1 0 1 0 , ˆλ7= 0 0 0 0 0−i 0i0 ,ˆλ8=1√ 3 1 0 0 0 1 0 0 0−2 , (6.126) 6.3. LIE ALGEBRAS 243 which form a basis for the real vector space of 3-by-3 tracele ss, hermitian matrices. They have been chosen and normalized so that tr (ˆλiˆλj) = 2δij, (6.127) by analogy with the properties of the Pauli matrices. Notice thatˆλ3andˆλ8 commute with each other, and that this will be true in any repr esentation. The matrices t±=1 2(ˆλ1±iˆλ2), v±=1 2(ˆλ4±iˆλ5), u±=1 2(ˆλ6±iˆλ7). (6.128) have unit entries, rather like the step up and step down matri cesσ±= 1 2(ˆσ1±iˆσ2). Let us define Λ ito be abstract operators with the same commutation relations as ˆλi, and define T±=1 2(Λ1±iΛ2), V±=1 2(Λ4±iΛ5), U±=1 2(Λ6±iΛ7). (6.129) These are simultaneous eigenvectors of the commuting pair o f operators ad (Λ 3) and ad(Λ 8): ad(Λ 3)T±= [Λ 3,T±] =±2T±, ad (Λ 3)V±= [Λ 3,V±] =±V±, ad(Λ 3)U±= [Λ 3,U±] =∓U±, ad(Λ 8)T±= [Λ 8,T±] = 0 ad (Λ 8)V±= [Λ 8,V±] =±√ 3V±, ad(Λ 8)U±= [Λ 8,U±] =±√ 3U±, (6.130) Thus, in any representation, the T±,U±,V±, act as ladder operators, chang- ing the simultaneous eigenvalues of the commuting pair Λ 3, Λ8. Their eigen- values,λ3,λ8, are called the weights , and there will be a set of such weights 244 CHAPTER 6. LIE GROUPS for each possible representation. By using the ladder opera tors one can go from any weight in a representation to any other, but you cann ot get outside this set. The amount by which the ladder operators change the weights are called the roots orroot vectors , and the root diagram characterizes the Lie algebra. + −+ −T+ T−U UV Vλ8 λ33 2 −2 − 3 Figure 6.2: The root vectors of su(3). In a finite-dimensional representation there must be a highe st weight state |λ3,λ8/angbracketrightthat is killed by all three of U+,T+andV+. We can then obtain all other states in the representation by repeatedly acting on the highest weight state with U−,T−orV−and their products. Since there is usually more than one route by which we can step down from the highest w eight to another weight, the weight spaces may be degenerate —i.ethere may be more than one linearly independent state with the same eigen values of Λ 3 and Λ 8. Exactly what states are obtained, and with what multiplici ty, is not immediately obvious. We will therefore restrict ourselves to describing the outcome of this procedure without giving proofs. What we find is that the weights in a finite-dimensional repres entation of su(3) form a hexagonally symmetric “crystal” lying on a triang ular lattice, and the representations may be labelled by pairs of integers (zero allowed) p,qwhich give the length of the sides of the crystal. These repre sentations have dimension d=1 2(p+ 1)(q+ 1)(p+q+ 2). 6.3. LIE ALGEBRAS 245 3 3333λ8 5 2 −7−4−1 0 λ3 −4 −3 −2 −1 2 4 1 3 Figure 6.3: The weight diagram of the 24 dimensional irrep with p= 3, q= 1. The highest weight is shaded. Figure 6.3 shows the set of weights occurring in the represen tation of SU(3) withp= 3 andq= 1. Each circle represents a state, whose weight ( λ3,λ8) may be read off from the displayed axes. A double circle indica tes that there are two linearly independent vectors with the same weight. A count confirms that the number of independent weights, and hence the dimens ion of the representation, is 24. For SU(3) representations the degen eracy— i.e.the number of states with a given weight—increases by unity at ea ch “layer” until we reach a triangular inner core, all of whose weights h ave the same degeneracy. In particle physics applications representations are ofte n labelled by their dimension. The defining representation of SU(3) and its comp lex conjugate are denoted by 3 and ¯3, 246 CHAPTER 6. LIE GROUPS 3 33 3λ8 λ3−1 1 01 λ3λ8 −1 1 02 −1 −2 Figure 6.4: The weight diagrams of the irreps with p= 1,q= 0, andp= 0, q= 1, also known, respectively, as the 3and the 3. while the weight diagrams of the eight dimensional adjoint r epresention and the 10 have shape shown in figure 6.5. Figure 6.5: The irreps 8(the adjoint) and 10. Cartan algebras: roots and co-roots For a general simple Lie algebra we may play the same game. We fi rst find a maximal linearly independent set of commuting generators, hi. Thehiform a basis for the Cartan algebra ,h, whose dimension is the rank of the Lie algbera. We next find ladder operators by diagonalizing the “ ad” action of thehion the rest of the algebra. ad(hi)eα= [hi,eα] =αieα. (6.131) The simultaneous eigenvectors eαare the ladder operators that change the eigenvalues of the hi. The corresponding eigenvalues α, thought of as vectors with components αi, are the roots, or root vectors. The roots are therefore the weights of the adjoint representation. It is possible to put factors of “ i” 6.3. LIE ALGEBRAS 247 in the appropriate places so that the αiare real, and we will assume that this has been done. For example in su(3) we have already seen that αT= (2,0), αV= (1,√ 3),αU= (−1,√ 3). Here are the basic properties and ideas that emerge from this process: i) Sinceαi/angbracketlefteα,hj/angbracketright=/angbracketleftad(hi)eα,hj/angbracketright=−/angbracketlefteα,[hi,hj]/angbracketright= 0 we see that /angbracketlefthi,eα/angbracketright= 0. ii) Similarly, we see that ( αi+βi)/angbracketlefteα,eβ/angbracketright= 0, so the eαare orthogonal to one another unless α+β= 0. Since our Lie algebra is semisimple, and consequently the Killing form non-degenerate, we deduce th at ifαis a root, so is−α. iii) Since the Killing form is non-degenerate, yet the hiare orthogonal to all theeα, it must also be non-degenerate when restricted to the Carta n algebra. Thus the metric tensor, gij=/angbracketlefthi,hj/angbracketright, must be invertible with inversegij. We will use the notation α·βto represent αiβjgij. iv) Ifα,βare roots, then the Jacobi identity shows that [hi,[eα,eβ]] = (αi+βi)[eα,eβ], so if [eα,eβ] is non-zero then α+βis also a root, and [ eα,eβ]∝eα+β. v) It follows from iv), that [ eα,e−α] commutes with all the hi, and since h was assumed maximal, it must either be zero or a linear combin ation of thehi. A short calculation shows that /angbracketlefthi,[eα,e−α]/angbracketright=αi/angbracketlefteα,e−α/angbracketright, and, since/angbracketlefteα,e−α/angbracketrightdoes not vanish, [ eα,e−α] is non-zero. Thus [eα,e−α]∝2αi α2hi≡hα whereαi=gijαj, andhαobeys [hα,e±α] =±2e±α. Thehαare called the co-roots . vi) The importance of the co-roots stems from the observatio n that the triadhα,e±αobey the same commutation relations as ˆ σ3andσ±, and so form an su(2) subalgebra of g. In particular hα(being the analogue of 2J3) has only integer eigenvalues. For example in su(3) [T+,T−] =hT= Λ3, 248 CHAPTER 6. LIE GROUPS [V+,V−] =hV=1 2Λ3+√ 3 2Λ8, [U+,U−] =hU=−1 2Λ3+√ 3 2Λ8, and in the defining representation hT= 1 0 0 0−1 0 0 0 0  hV= 1 0 0 0 0 0 0 0−1  hU= 0 0 0 0 1 0 0 0−1 , have eigenvalues±1. vii) Since ad (hα)eβ= [hα,eβ] =2α·β α2eβ, we conclude that 2 α·β/α2must be an integer for any pair of roots α, β. viii) Finally, there can only be one eαfor each root α. If not, and there were an independent e/prime α, we could take linear combinations so that e−α ande/prime αare Killing orthogonal, and hence [ e−α,e/prime α] =αihi/angbracketlefte−α,e/prime α/angbracketright= 0. Thus ad (e−α)e/prime α= 0, ande/prime αis killed by the step-down operator. It would therefore be the lowest weight in some su(2) representation. At the same time, however, ad ( hα)e/prime α= 2e/prime α, and we know that the lowest weight in any spin Jrepresentation cannot have positive eigenvalue. The conditions that 2α·β α2∈Z for any pair of roots tightly constrains the possible root sy stems, and is the key to Cartan and Killing’s classification of the semisimple Lie algebras. For example the angle θbetween any pair of roots obeys cos2θ=n/4 soθcan take only the values 0◦,30◦,45◦,60◦,90◦,120◦,135◦,150◦, or 180◦. 6.3. LIE ALGEBRAS 249 These constraints lead to a complete classification of possi ble root systems into the infinite families An, n= 1,2,···. sl(n+ 1,C), Bn, n= 2,3,···. so(2n+ 1,C), Cn, n= 3,3,···. sp(2n,C), Dn, n= 4,5,···. so(2n,C), together with the root systems G2,F4,E6,E7, andE8of the exceptional algebras. The latter do not correspond to any of the classica l matrix groups. For example G2is the root system of g2, the Lie algebra of the group G 2of automorphisms of the octonions . This group is also the subgroup of SL(7) preserving the general totally antisymmetric trilinear fo rm. The restrictions on n’s are to avoid repeats arising from “accidental” isomorphisms. If we allow n= 1,2,3, in each series, then C1=D1=A1. This corresponds to sp(2,C)∼=so(3,C)∼=sl(2,C). Similarly D2=A1+A1, corresponding to isomorphism SO(4) ∼=SU(2)×SU(2)/Z2, whileC2=B2 implies that, locally, the compact Sp(2) ∼=SO(5). Finally D3=A3implies that SU(4)/Z2∼=SO(6). 6.3.4 Product Representations Given two representations Λ(1) iand Λ(2) iofg, we can form a new representa- tion that exponentiates to the tensor product of the corresp onding represen- tations of the group G. Motivated by the result of exercise 5.13: exp(A⊗In+Im⊗B) = exp(A)⊗exp(B) (6.132) we set Λ(1⊗2) i= Λ(1) i⊗I(2)+I(1)⊗Λ(2) i. (6.133) Then [Λ(1⊗2) i,Λ(1⊗2) j] = ([Λ(1) i⊗I(2)+I(1)⊗Λ(2) i),(Λ(1) j⊗I(2)+I(1)⊗Λ(2) j)] = [Λ(1) i,Λ(1) j]⊗I(2)+ [Λ(1) i,I(1)]⊗Λ(2) j +Λ(1) i⊗[I(2),Λ(2) j] +I(1)⊗[Λ(2) i,Λ(2) j] = [Λ(1) i,Λ(1) j]⊗I(2)+I(1)⊗[Λ(2) i,Λ(2) j], (6.134) 250 CHAPTER 6. LIE GROUPS showing that the Λ(1⊗2) ialso obey the Lie algebra. This process of combining representations is analogous to t he addition of angular momentum in quantum mechanics. Perhaps more prec isely, the addition of angular momentum is an example of this general co nstruction. If representation Λ(1) ihas weights m(1) i,i.e.h(1) i|m(1)/angbracketright=m(1) i|m(1)/angbracketright, and Λ(2) i has weights m(2) i, then, writing|m(1),m(2)/angbracketrightfor|m(1)/angbracketright⊗|m(2)/angbracketright, we have h(1⊗2) i|m(1),m(2)/angbracketright= (h(1) i⊗1 + 1⊗h(2) i)|m(1),m(2)/angbracketright = (m(1) i+m(2) i)|m(1),m(2)/angbracketright (6.135) so the weights appearing in the representation Λ(1⊗2) iarem(1) i+m(2) i. The new representation is usually decomposible. We are fami liar with this decomposition for angular momentum where, if j >j/prime, j⊗j/prime= (j+j/prime)⊕(j+j/prime−1)⊕···(j−j/prime). (6.136) This can be understood from adding weights. For example cons ider adding the weights of j= 1/2, which are m=±1/2 to those of j= 1, which are m=−1,0,1. We getm=−3/2,−1/2 (twice) +1 /2 (twice) and m= 3/2. These decompose as shown in figure 6.6. = Figure 6.6: The weights for 1/2⊗1 = 3/2⊕1/2. The rules for decomposing products in other groups are more c ompli- cated than for SU(2), but can be obtained from weight diagram s in the same manner. In SU(3), we have, for example 3⊗¯3 = 1⊕8, 3⊗8 = 3⊕¯6⊕15, 8⊗8 = 1⊕8⊕8⊕10⊕10⊕27. (6.137) To illustrate the first of these we show, in figure 6.7 the addit ion of the weights in ¯3 ) to each of the weights in the 3. 6.3. LIE ALGEBRAS 251 = Figure 6.7: Adding the weights of 3and¯3. The resultant weights decompose (uniquely) into the weight diagrams for the 8 together with a singlet. 6.3.5 Sub-algebras and branching rules As with finite groups, a representation that is irreducible u nder the full Lie group or algebra will in general become reducible when restr icted to a sub- group or sub-algebra. The pattern of the decomposition is ag ain called a branching rule . Here we provide some examples to illustrate the ideas. The three operators V±andhV=1 2Λ3+√ 3 2Λ8ofsu(3) form a Lie sub- algebra that is isomorphic to su(2) under the map that takes them to σ± andσ3respectively. When restricted to this sub-algebra, the 8 di mensional representation of su(3) becomes reducible, decomposing as 8 = 3⊕2⊕2⊕1, (6.138) where the 3, 2 and 1 are the j= 1,1 2and 0 representations of su(2). We can visualize this decomposition coming about by first pro jecting the (λ3,λ8) weights to the “ m” of the|j,m/angbracketrightlabelling of su(2) as m=1 4λ3+√ 3 4λ8 (6.139) and then stripping off the su(2) irreps as we did when decomposing product representions. 252 CHAPTER 6. LIE GROUPS m=1 m=1/2 m=0 m=−1/2 m=−1 Figure 6.8: Projection of the su(3)weights onto su(2), and the decomposition 8 = 3⊕2⊕2⊕1. This branching pattern occurs in the strong interactions wh ere the mass of the strange quark sbeing much larger than that of the light quarks uand dcauses the octet of pseudo-scalar mesons, which would all ha ve the same mass if SU(3) flavour symmetry was exact, to decompose into th e triplet of pionsπ+,π0andπ−, the pairK+andK0, their antiparticles K−and¯K0, and the singlet η. There are obviously other su(2) sub-algebras consisting of {T±,hT}and {U±,hU}, each giving rise to similar decompositions. These sub-alg ebras, and a continuous infinity of related ones, are obtained from t he{V±,hV} algebra by conjugation by elements of SU(3). Another, unrelated, su(2) sub-algebra consists of σ+/similarequal√ 2(U++T+), σ−/similarequal√ 2(U−+T−), σ3/similarequal2hV= (Λ 3+√ 3Λ8). (6.140) The factor of two between the assignment σ3/similarequalhVof our previous example and the present assignment σ3/similarequal2hVhas a non-trivial effect on the branching rules. Under restriction to this new subalgebra, the 8 of su(3) decomposes as 8 = 5⊕3 (6.141) 6.4. FURTHER EXERCISES AND PROBLEMS 253 m=2 m=1 m=0 m=−1 m=−2 Figure 6.9: The projection and decomposition for 8 = 5⊕3. where the 5 and 3 are the j= 2 andj= 1 representations of su(2). A clue to the origin and significance of this sub-algebra is found by noting that the 3 and ¯3 representations of su(3) both remain irreducible, but project to the samej= 1 representation of su(2). Interpreting this j= 1 representation as the defining vector representation of so(3) suggests (correctly) that our newsu(2) sub-algebra is the Lie algebra of the SO(3) subgroup of SU (3) consisting of SU(3) matrices with real entries. 6.4 Further Exercises and Problems Exercise 6.25 :Campbell-Baker-Hausdorff Formulae . Here are some useful formula for working with exponentials of matrices that do no t commute with each other. a) LetXandXbe matrices. Show that etXYe−tX=Y+t[X,Y] +1 2t2[X,[X,Y]] +···, the terms on the right being the series expansion of exp[ad( tX)]Y. b) LetXandδXbe matrices. Show that e−XeX+δX= 1 +/integraldisplay1 0e−tXδXetXdt+O/bracketleftbig (δX)2/bracketrightbig = 1 +δX−1 2[X,δX ] +1 3![X,[X,δX ]] +···+O/bracketleftbig (δX)2/bracketrightbig = 1 +/parenleftBigg 1−ead(X) ad(X)/parenrightBigg δX+O/bracketleftbig (δX)2/bracketrightbig (6.142) c) By expanding out the exponentials, show that eXeY=eX+Y+1 2[X,Y]+higher, 254 CHAPTER 6. LIE GROUPS where “higher” means terms higher order in X,Y. The next two terms are, in fact,1 12[X,[X,Y]]+1 12[Y,[Y,X]]. You will find the general formula in part d). d) By using the formula from part b), show that that eXeYcan be written aseZ, where Z=X+/integraldisplay1 0g(ead(X)ead(tY))Y dt. Here g(z)≡lnz 1−1/z has a power series expansion g(z) = 1 +1 2(z−1) +1 6(z−1)2+1 12(z−1)3+···, which is convergent for |z|<1. Show that g(ead(X)ead(tY)) can be ex- panded as a double power series in ad( X) and ad(tY), provided Xand Yare small enough. This ad( X), ad(tY) expansion allows us to evaluate the product of two matrix exponentials as a third matrix expo nential provided we know their commutator algebra. Exercise 6.26 : SU(2) Disentangling theorems : Almost any 2×2 matrix can be factored (Gaussian decomposition) as /parenleftbigga b c d/parenrightbigg =/parenleftbigg1α 0 1/parenrightbigg/parenleftbiggλ0 0µ/parenrightbigg/parenleftbigg1 0 β1/parenrightbigg . Use this trick to work the following problems: a) Show that exp/braceleftbiggθ 2(eiφˆσ+−e−iφˆσ−)/bracerightbigg = exp(αˆσ+)exp(λˆσ3)exp(βˆσ−), where ˆσ±= (ˆσ1±iˆσ2)/2, and α=eiφtanθ/2, λ=−ln cosθ/2, β=−e−iφtanθ/2. b) Use the fact that the spin-1 2representation of SU(2) is faithful, to show that exp/braceleftbiggθ 2(eiφˆJ+−e−iφˆJ−)/bracerightbigg = exp(αˆJ+)exp(2λˆJ3)exp(βˆJ−), 6.4. FURTHER EXERCISES AND PROBLEMS 255 where ˆJ±=ˆJ1±iˆJ2. Take care, the reasoning here is subtle! Notice that the series expansion of exponentials of ˆ σ±truncates after the second term, but the same is nottrue of the expansion of exponentials of the ˆJ±. You need to explain why the formula continues to hold in the ab sence of this truncation. Exercise 6.27 :Invariant tensors for SU(3). Let λibe the Gell-Mann lambda matrices. The totally antisymmetric structure constants, fijk, and a set of totally symmetric constants dijkare defined by fijk=1 2tr (λi[λj,λk]), dijk=1 2tr (λi{λj,λk}). LetD8 ij(g) be the matrices representing SU(3) in “8” — the eight-dimen sional adjoint representation. a) Show that fijk=D8 il(g)D8 jm(g)D8 kn(g)flmn, dijk=D8 il(g)D8 jm(g)D8 kn(g)dlmn, and sofijkanddijkareinvariant tensors in the same sense that δijand /epsilon1i1...inare invariant tensors for SO( n). b) Letwi=fijkujvk. Show that if ui→D8 ij(g)ukandvi→D8 ij(g)vk, then wi→D8 ij(g)wk. Similarly for wi=dijkujvk. (Hint: show first that theD8matrices are real and orthogonal.) Deduce that fijkanddijkare Clebsh-Gordan coefficients for the 8⊕8 part of the decomposition 8⊗8 = 1⊕8⊕8⊕10⊕10⊕27. c) Similarly show that δαβand the lambda matrices ( λi)αβcan be regarded as Clebsch-Gordan coefficients for the decomposition ¯3⊗3 = 1⊕8. d) Use the graphical method of plotting weights and peeling o ff irreps to obtain the tensor product decomposition in part b). 256 CHAPTER 6. LIE GROUPS Chapter 7 The Geometry of Fibre Bundles In earlier chapters we have used the language of bundles and c onnections, but in a relatively casual manner. We deferred proper mathemati cal definitions until now, because, for the applications we meet in physics, it helps to first have acquired an understanding of the geometry of Lie groups . 7.1 Fibre Bundles We begin with a formal definition of a bundle and then illustra te the defini- tion with examples from quantum mechanics. These allow us to appreciate the physics that the definition is designed to capture. 7.1.1 Definitions A smooth bundle is a triple ( E,π,M ) whereEandMare manifolds, and π:E→Mis a smooth map. The manifold Eis called the total space ,M is the base space andπthe projection map. The inverse image π−1(x) of a point inM(i.e.the set of points in Ethat map to xinM), is the fibreover x. We usually require that all fibres be diffeomorphic to some fixe d manifold F. The bundle is then a fibre bundle , andFis “the fibre” of the bundle. In a similar vein, we sometimes also refer to the total space Eas “the bundle.” Examples of possible fibres are vector spaces (in which case w e have a vector bundle ), spheres (in which case we have a sphere bundle), and Lie gro ups. When the fibre is a Lie group we speak of a principal bundle . A principal 257 258 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES bundle can be thought of the parent of various associated bundles , which are constructed by allowing the Lie group to act on a fibre. A bundl e whose fibre is a one dimensional vector space is called a line bundle . The simplest example of a fibre bundle consists of setting Eequal to the Cartesian product M×Fof the base space and the fibre. In this case the projection just “forgets” the point f∈F, and soπ: (x,f)/mapsto→x. A more interesting example can be constructed by taking Mto be the circleS1, andFas the one-dimensionsional interval I= [−1,1]. We can assemble these ingredients to make Einto a M¨ obius strip . We do this by gluing the copy of Ioverθ= 2πto that over θ= 0 with a half twist so that the end−1∈[−1,1] is attached to +1, and vice versa. +1 φ−1 −1+1 0 2π0 SE 1π Figure 7.1: M¨ obius strip bundle, together with a section φ. A bundle that is a product E=M×F, is said to be trivial . The M¨ obius strip is not a Cartesian product, and is said to be a twisted bundle. The M¨ obius strip is, however, locally trivial in that for each x∈Mthere is an open retractable neighbourhood U⊂Mofxin whichElooks like a product U×F. We will assume that all our bundles are locally trivial in th is sense. If {Ui}is a cover of M(i.e.ifM=/uniontextUi) by such retractable neighbourhoods, andFis a fixed fibre, then a bundle can be assembled out of the collec tion ofUi×Fproduct bundles by giving gluing rules that identify points on the fibre overx∈Uiin the product Ui×Fwith points in the fibre over x∈Uj inUj×Ffor eachx∈Ui∩Uj. These identifications are made by means of invertible maps ϕUiUj(x) :F→Fthat are defined for each xin the overlap 7.2. PHYSICS EXAMPLES 259 Ui∩Uj. TheϕUiUjare known as transition functions . They must satisfy the consistency conditions ϕUiUi(x) = Identity , ϕUiUj(x) =φ−1 UjUi(x) ϕUiUj(x)ϕUjUk(x) =ϕUiUk(x), x∈Ui∩Uj∩Uk/negationslash=∅. (7.1) Asection of a fibre bundle ( E,π,M ) is a smooth map φ:M→Esuch thatφ(x) lies in the fibre π−1(x) overx. Thusπ◦φ= Identity. When the total space Eis a product M×Fthisφis simply a function φ:M→F. When the bundle is twisted, as is the M¨ obius strip, then the s ection is no longer a function as it takes no unique value at the points xabove which the fibres are being glued together. Observe that in the M¨ obi us strip the half-twist forces the section φ(x) to pass through 0 ∈[−1,1]. The M¨ obius bundle therefore has no nowhere-zero globally defined secti ons. Many twisted bundles have no globally defined sections at all. 7.2 Physics Examples We now provide three applications where the bundle concept a ppears in quantum mechanics. The first two illustrations are re-expre ssions of well- known physics. The third, the geometric approach to quantiz ation, is perhaps less familiar. 7.2.1 Landau levels Consider the Schr¨ odinger eigenvalue problem −1 2m/parenleftbigg∂2ψ ∂x2+∂2ψ ∂y2/parenrightbigg =Eψ (7.2) for a particle moving on a flat two-dimensional torus. We thin k of the torus as aLx×Lyrectangle with the understanding that as a particle disappe ars through the right-hand boundary it immediately re-appears at the point with the sameyco-ordinate on the left-hand boundary; similarly for the up per and lower boundaries. In quantum mechanics we implement the se rules by imposing periodic boundary conditions on the wave function : ψ(0,y) =ψ(Lx,y)ψ(x,0) =ψ(x,Ly). (7.3) 260 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES These conditions make the wavefunction a well-defined and co ntinuous func- tion on the torus, in the sense that after pasting the edges of the rectangle together to make a real toroidal surface the function has no j umps, and each point on the surface assigns a unique value to ψ. The wavefunction is a section of an untwisted line bundle with the torus as its base -space, the fi- bre over (x,y) being the one-dimensional complex vector space Cin which ψ(x,y) takes its value. Now try to carry out the same program for a particle of charge emoving in a uniform magnetic field Bperpendicular to the x−yplane. The Schr¨ odinger equation becomes −1 2m/parenleftbigg∂ ∂x−ieAx/parenrightbigg2 ψ−1 2m/parenleftbigg∂ ∂y−ieAy/parenrightbigg2 ψ=Eψ, (7.4) where (Ax,Ay) is the vector potential. We at once meet a problem. Although the magnetic field is constant, the vector potential cannot b e chosen to be constant — or even periodic. In the Landau gauge , for example, where we set Ax= 0, the remaining component becomes Ay=Bx. This means that as the particle moves out of the right-hand edge of the rectangle re presenting the torus we must perform a gauge transformation that prepares i t for motion in the (Ax,Ay) field it will encounter when it reappears at the left. If (7.4 ) holds, then it continues to hold after the simultaneous chan ge ψ(x,y)→e−ieBL xyψ(x,y) −ieAy→ −ieAy+e−iBL xy∂ ∂ye+ieBL xy=−ie(Ay−BLx).(7.5) At the right-hand boundary x=Lxthis gauge transformation resets the vector potential Ayback to its value at the left-hand boundary. Accordingly, we modify the boundary conditions to ψ(0,y) =e−ieBL xyψ(Lx,y), ψ (x,0) =ψ(x,Ly). (7.6) The new boundary conditions make the wavefunction into a sec tion1of a —it twisted line bundle over the torus. The fibre is again the one- dimensional complex vector space C. 1That the wave “function” is no longer a function should not be disturbing. Schr¨ odinger’s ψis never really a function of space-time. Seen from a frame moving at velocityv,ψ(x,t) acquires factor of exp( −imvx−mv2t/2), and this is no way for a self- respecting function of xandtto behave. 7.2. PHYSICS EXAMPLES 261 We have already met the language in which the gauge field −ieAµis a called connection on the bundle, and the associated ieBfield is the curvature . We will explain how connections fit into the formal bundle lan guage in section 7.3. The twisting of the boundary conditions by the gauge transfo rmation seems innocent, but within it lurks an important constraint related to the consistency conditions in (7.1). We can find the value of ψ(Lx,Ly) from that ofψ(0,0) by using the relations in (7.6) in the order ψ(0,0)→ψ(0,Ly)→ ψ(Lx,Ly), or in the order ψ(0,0)→ψ(Lx,0)→ψ(Lx,Ly). Since we must obtain the same ψ(Lx,Ly) whichever route we use, we need to satisfy the condition eieBL xLy= 1. (7.7) This tells us that the Schr¨ odinger problem makes sense only when the mag- netic fluxBLxLythrough the torus obeys eBLxLy= 2πN (7.8) for some integer N. We cannot continuously vary the flux through a fi- nite torus. This means that if we introduce torus boundary co nditions as a mathematical convenience in a calculation, then physical e ffects may depend discontinuously on the field. The integer Ncounts the number of times the phase of the wavefunction is twisted as we travel from x=Lx,y= 0 tox=Lx,y=Lygluing the right-hand edge wavefunction to back to the left-hand edge w avefunction. This twisting number is a topological invariant. We have met this invariant before, in section 4.6. It is the first Chern number of the wavefunction bundle. If we permit Bto become position without altering the total twist N, then quantities such as energies and expectation values can change smoothly withB. IfNis allowed to change, however, the these quantities may jump discontinuously. The energy E=Ensolutions to (7.4) with boundary conditions (7.6) are given by Ψn,k(x,y) =∞/summationdisplay p=−∞ψn/parenleftbigg x−k B−pLx/parenrightbigg ei(eBpL x+k)y. (7.9) Hereψn(x) is a harmonic-oscillator wavefunction obeying −1 2md2ψn dx2+1 2mω2ψn=Enψn, (7.10) 262 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES withω=eB/m the classical cyclotron frequency, and En=ω(n+1/2). The parameterktakes the values 2 πq/Lyforqan integer. At each energy Enwe obtainNindependent eigenfunctions as qruns from 1 to eBLxLy/2π. These N-fold degenerate states are the Landau levels . The degeneracy, being of necessity an integer, provides yet another explanation for why the flux must be quantized. 7.2.2 The Berry connection Suppose we are in possession of a quantum-mechanical hamilt onian ˆH(ξ) de- pending on some parameters ξ= (ξ1,ξ2,...)∈M, and know the eigenstates |n;ξ/angbracketrightthat obey ˆH(ξ)|n;ξ/angbracketright=En(ξ)|n;ξ/angbracketright. (7.11) If, for fixed n, we can find a smooth family of eigenstates |n;ξ/angbracketright, one for everyξin the parameter space M, we have a vector bundle over the space M. The fibre above ξis the one-dimensional vector space spanned by |n;ξ/angbracketright. This bundle is a sub-bundle of the product bundle M×HwhereHis the Hilbert space on which ˆHacts. Although the larger bundle is not twisted, the sub-bundle may be. It may also not exist: if the state |n;ξ/angbracketrightbecome degenerate with another state |m;ξ/angbracketrightat some value of ξ, then both states can vary discontinuously with the parameters, and we wish to exclude this possibility. In the previous paragraph we considered the evolution of the eigenstates of a time-independent Hamiltonian as we varied its paramete rs. Another, more physical, evolution is given by solving the time-dependent Schr¨ odinger equation i∂t|ψ(t)/angbracketright=ˆH(ξ(t))|ψ(t)/angbracketright (7.12) so as to follow the evolution of a state |ψ(t)/angbracketrightas the parameters are slowly var- ied. If the initial state |ψ(0)/angbracketrightcoincides with with the eigenstate |0,ξ(0)/angbracketright, and if the time evolution of the parameters is slow enough, then |ψ/angbracketrightis expected to remain close to the corresponding eigenstate |0;ξ(t)/angbracketrightof the time-independent Schr¨ odinger equation for the hamiltonian ˆH(ξ(t)). To determine exactly how “close” it stays, insert the expansion |ψ(t)/angbracketright=/summationdisplay nan(t)|n;ξ(t)/angbracketrightexp/braceleftbigg −i/integraldisplayt 0E0(ξ(t))dt/bracerightbigg . (7.13) 7.2. PHYSICS EXAMPLES 263 into (7.12) and take the inner-product with |m;ξ/angbracketright. Form/negationslash= 0, we expect that the overlap/angbracketleftm;ξ|ψ(t)/angbracketrightwill be small and of order O(∂ξ/∂t ). Assuming that this is so, we read off that ˙a0+a0/angbracketleft0;ξ|∂µ|0;ξ/angbracketright∂ξµ ∂t= 0,(m= 0) (7.14) am=ia0/angbracketleftm;ξ|∂µ|0;ξ/angbracketright Em−E0∂ξµ ∂t,(m/negationslash= 0) (7.15) up to first-order accuracy in time derivatives of the |n;ξ(t)/angbracketright. Hence |ψ(t)/angbracketright=eiγBerry(t)/braceleftBigg |0;ξ/angbracketright+i/summationdisplay m/negationslash=0|m;ξ/angbracketright/angbracketleftm;ξ|∂µ|0;ξ/angbracketright Em−E0∂ξµ ∂t+.../bracerightBigg e−iRt 0E0(t)dt, (7.16) where the dots refer to terms of higher order in time derivati ves. Equation (7.16) constitutes the first two terms in a systemat icadiabatic series expansion . The factor a0(t) = exp{iγBerry(t)}is the solution of the differential equation (7.14). The angle γBerryis known as Berry’s phase after the British mathematical physicist Michael Berry. It is nee ded to take up the slack between the arbitrary ξ-dependent phase choice at our disposal when defining the|0;ξ/angbracketright, and the specific phase selected by the Schr¨ odinger equatio n as it evolves the state |ψ(t)/angbracketright. Berry’s phase is also called the geometric phase because it depends only on the Hillbert-space geometry of th e family of states |0;ξ/angbracketright, and not on their energies. We can write γBerry(t) =i/integraldisplayt 0/angbracketleft0;ξ|∂µ|0;ξ/angbracketright∂ξµ ∂tdt (7.17) and regard the one-form ABerrydef=/angbracketleft0;ξ|∂µ|0;ξ/angbracketrightdξµ=/angbracketleft0;ξ|d|0;ξ/angbracketright (7.18) as a connection on the bundle of states over the space of param eters. The equation ˙ξµ/parenleftbigg∂ ∂ξµ+ABerry,µ/parenrightbigg ψ= 0 (7.19) then identifies the Schr¨ odinger time evolution with parall el transport. It seems reasonable to refer to this particular parallel trans port as “Berry trans- port.” 264 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES In order for corrections to the approximation |ψ(t)/angbracketright≈(phase)|0;ξ(t)/angbracketrightto remain small, we need the denominator ( Em−E0) to remain large when compared to its numerator. The state that we are following mu st therefore never become degenerate with any other state. Monople bundle Consider, for example a spin-1 /2 particle in a magnetic field. If the field points in direction n, the Hamiltonian is ˆH(n) =µ|B|ˆσ·n (7.20) There are are two eigenstates with energy E±=±µ|B|. Let is focus on the eigenstate|ψ+/angbracketrightcorresponding to E+. For each nwe can obtain an E+ eigenstate by applying the projection operator ˆP=1 2(I+n·ˆσ) =1 2/parenleftbigg 1 +nznx−iny nx+iny1−nz/parenrightbigg (7.21) to almost any vector, and then multiplying by a real normaliz ation constant N. Applying ˆPto a “spin-up” state, for example gives =N1 2(I+n·ˆσ)/parenleftbigg 1 0/parenrightbigg =/parenleftbigg cosθ/2 eiφsinθ/2/parenrightbigg . (7.22) Hereθandφare spherical polar angles on S2that specify the direction of n. Although the bundle of E=E+eigenstates is globally defined, the family of states|ψ(1) +(n)/angbracketrightthat we have obtained, and would like to use as base for the fibre over n, becomes singular when nis in the vicinity of the south pole θ=π. This is because the factor eiφis multivalued at the south pole. There is no problem at the north pole because the ambiguous phase eiφmultiples sinθ/2, which is zero there. Near the south pole, however, we can project from a “spin-dow n” state to find. |ψ(2) +(n)/angbracketright=N1 2(I+n·ˆσ)/parenleftbigg 0 1/parenrightbigg =/parenleftbigg e−iφcosθ/2 sinθ/2/parenrightbigg . (7.23) This family of eigenstates is smooth near the south pole, but is ill-defined at the north pole. As in section 4.6, we are compelled to cover th e sphereS2 7.2. PHYSICS EXAMPLES 265 by two caps D+andD−, and use|ψ(1) +/angbracketrightinD+and|ψ(2) +/angbracketrightinD−. The two families are related by |ψ(1) +(n)/angbracketright=eiφ|ψ(2) +(n)/angbracketright (7.24) in the cingular overlap region D+∩D−. Hereeiφis the transition function that glues the two families of eigenstates together. The Berry connections are A(1) +=/angbracketleftψ(1) +|d|ψ(1) +/angbracketright=i 2(cosθ−1)dφ A(2) +=/angbracketleftψ(2) +|d|ψ(2) +/angbracketright=i 2(cosθ+ 1)dφ. (7.25) In their common domain of definition, they are related by a gau ge transfor- mation A(2) +=A(1) ++idφ. (7.26) The curvature of either connection is dA=−i 2sinθdθdφ =−i 2d(Area). (7.27) The curvature being the area two-form tells us that when we sl owly change the direction of Band bring it back to its original orientation the spin state will, in addition to the dynamical phase exp{−iE+t}, have accumulated a phase equal to (minus) one-half of the area enclosed by the tr ajectory of n onS2. The two-form field dAcan be though of as the flux of a magnetic monople residing at the centre of the sphere. The bundle of on e-dimensional vector spaces span[ |ψ+(n)/angbracketright] overS2is therefore called the monople bundle . 7.2.3 Quantization In this section we provide a short introduction to geometric quantization . This idea, due largely to Kirilov, Kostant and Souriau, exte nds the famil- iar technique of canonical quantization to phase spaces wit h more structure than that of the harmonic oscillator. We illustrate the form alism by quan- tizing spin, and show how the resulting Hilbert space provid es an example of the Borel-Weil-Bott construction of the representations o f a semi-simple Lie group as spaces of sections of holomorphic line bundles. 266 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES Prequantization The passage from classical mechanics to quantum mechanics i nvolves re- placing the classical variables by operators in such a way th at the classical Poisson-bracket algebra is mirrored by the operator commut ator algebra. In general, this process of quantization is not possible without making some compromises. It is, however, usually possible to pre-quantize a phase-space with its associated Poisson algebra. LetMbe a 2n-dimensional classical phase-space with its closed symple c- tic formω. Classically a function f:M→Rgive rise to a Hamiltonian vector field vfviaHamilton’s equations df=−ivfω. (7.28) We saw in section 2.4.2 that the closure condition dω= 0 ensures that that the Poisson bracket {f,g}=vfg=ω(vf,vg) (7.29) obeys [vf,vg] =v{f,g,}. (7.30) Now suppose that the cohomology class of (2 π/planckover2pi1)−1ωinH2(M,R) has the property that its integrals over cycles in H2(M,Z) are integers. Then (it can be shown) there exists a line bundle LoverMwith curvature F=−i/planckover2pi1−1ω. If we locally write ω=dη, whereη=ηµdxµ, then the connection one-form isA=−i/planckover2pi1−1ηand the covariant derivative ∇v≡vµ(∂µ−i/planckover2pi1−1ηµ), (7.31) acts on sections of the Line bundle. The corresponding curva ture is F(u,v)=[∇u,∇v]−∇ [u,v]=−i/planckover2pi1−1ω(u,v). (7.32) We define a pre-quantized operator /hatwideρ(f) that acting on sections Ψ( x) of the line bundle corresponds to the classical function f: /hatwideρ(f)def=−i/planckover2pi1∇vf+f. (7.33) For hamiltonian vector fields vfandvgwe have [/planckover2pi1∇vf+if,∇vg] = /planckover2pi1∇[vf,vg]−iω(vf,vg) +i[f,∇vg] =/planckover2pi1∇[vf,vg]−i(ivfω+df)(vg) =/planckover2pi1∇[vf,vg], (7.34) 7.2. PHYSICS EXAMPLES 267 and so [−i/planckover2pi1∇vf+f,−i/planckover2pi1∇vg+g] =−/planckover2pi12∇[vf,vg]−i/planckover2pi1vfg =−i/planckover2pi1(−i/planckover2pi1∇[vf,vg]+{f,g}) =−i/planckover2pi1(−i/planckover2pi1∇v{f,g}+{f,g}).(7.35) Equation (7.35) is Dirac’s quantization rule: i[/hatwideρ(f),/hatwideρ(g)] =/planckover2pi1/hatwideρ({f,g}). (7.36) The process of quantization is completed, when possible, by defining a polarization . This is a restriction on the variables that we allow the wave - functions to depend on. For example, if there is a global set o f Darboux co-ordinates p,qwe may demand that the wavefunction depend only on q, or only on the combination p+iq. Such a restriction is necessary so that the representation f/mapsto→/hatwideρ(f) isirreducible. Since globally defined Darboux co-ordinates do not usually exist, this step is the hard part of quantization. The precise definition of a polarized section is rather compl icated. We can only sketch it here, but give a concrete example in the nex t section. At each pointx∈Mthe symplectic form defines a skew bilinear form. We seek a Lagrangian subspace of Vx⊂TMpfor this form. A Lagrangian subspace is one such that Vx=V⊥ x. For example, if ω=dp1∧dq1+dp2∧dq2, (7.37) then the space spanned by the ∂q’s is Lagrangian, as is the space spanned by the∂p’s. We allow the coefficients of the vectors in Vxto be complex numbers. The vectors fields spanning the Vx’s form a distribution. We require it to be integrable, so that the Vxare the tangent spaces to a global foliation of M. A section Ψ of the Line bundle is polarized if∇ξΨ = 0 for all ¯ξ∈Vx. We define an inner product on the space of polarized sections b y using the Liouville measure ωn/n! on the phase space. The quantum Hilbert space then consists of finite-norm polarized sections of L. Only classical functions that give rise to polarization-compatible vector fields wil l have their Poisson- bracket algebra coincide with the quantum commutator algeb ra. Quantizing spin To illustrate these ideas, we quantize spin. The classical m echanics of spin was discussed in section 2.4.2. There we showed that the appr opriate phase 268 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES space is the 2-sphere equipped with a symplectic form propor tional to the area form. Here we must be specific about the constant of propo rtionality. We choose units in which /planckover2pi1→1, and take ω=jd(Area). The integrality of ω/2πrequires that jbe an integer or half integer. We will assume that jis positive. We parametrize the 2-sphere with complex sterographic co-o rdinatesz, zwhich are constructed similarly to those in section 3.4.3. T his choice will allow us to impose a natural complex polarization on the wave functions. In contrast to section 3.4.3, however, it is here convenient to make the point z= 0 correspond to the south pole, so the polar co-ordinates θ,φ, on the sphere are related to z,zvia cosθ=|z|2−1 |z|2+ 1, eiφsinθ=2z |z|2+ 1, e−iφsinθ=2z |z|2+ 1. (7.38) In terms of the z,zco-ordinates ω=2ij (1 +|z|2)2dz∧dz. (7.39) As long as we avoid the north pole where z=∞, we can write ω=d/braceleftbigg ijzdz−zdz 1 +|z|2/bracerightbigg =dη, (7.40) and so the local connection form has components proportiona l to ηz=−ijz |z|2+ 1, ηz=ijz |z|2+ 1. (7.41) The covariant derivatives are therefore ∇z=∂ ∂z−jz |z|2+ 1,∇z=∂ ∂z+jz |z|2+ 1. (7.42) We impose the polarization condition that ∇zΨ = 0. This condition requires the allowed sections to be of the form Ψ(z,z) = (1 +|z|2)−jψ(z), (7.43) 7.2. PHYSICS EXAMPLES 269 whereψdepends only on z. It is natural to combine the (1+ |z|2)−jprefactor with the Liouville measure so that the inner product becomes /angbracketleftψ|χ/angbracketright=2j+ 1 2πi/integraldisplay Cdz∧dz (1 +|z|2)2j+2ψ(z)χ(z). (7.44) The normalizable wavefunctions are then polynomials in zof degree less than or equal to 2 j, and a complete orthonormal set is given by ψm(z) =/radicalBigg 2j! (j−m)!(j+m)!zj+m,−j≤m≤j. (7.45) We desire to find the quantum operators /hatwideρ(Ji) corresponding to the com- ponents J1=jsinθcosφ, J 2=jsinθsinφ, J 3=jcosθ, (7.46) of a classical spin Jof magnitude j, and also to the ladder-operator compo- nentsJ±=J1±iJ2. In our complex co-ordinates these functions become J3=j|z|2−1 |z|2+ 1, J+=j2z |z|2+ 1, J−=j2z |z|2+ 1. (7.47) Hamilton’s equations read ˙z=i(1 +|z|2)2 2j∂H ∂z, ˙z=−i(1 +|z|2)2 2j∂H ∂z, (7.48) and the Hamiltonian vector fields corresponding to the class ical phase space functionsJ3,J+andJ−are vJ3=iz∂z−iz∂z, vJ+=−iz2∂z−i∂z, vJ−=i∂z+iz2∂z. (7.49) 270 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES Using the recipe (7.33) for /hatwideρ(H) from the previous section, and the fact that∇zΨ = 0, we find, for example, that /hatwideρ(J+)(1 +|z|2)−jψ(z) =/bracketleftbigg −z2/parenleftbigg∂ ∂z−jz (1 +|z|2)/parenrightbigg +2jz (1 +|z|2)/bracketrightbigg (1 +|z|2)−jψ(z), = (1 +|z|2)−j/bracketleftbigg −z2∂ ∂z+ 2jz/bracketrightbigg ψ (7.50) It is natural to define operators /hatwideJi= (1 +|z|2)j/hatwideρ(Ji)(1 +|z|2)−j(7.51) that act only on the z-polynomial part ψ(z) of the section Ψ( z,z). We then have /hatwideJ+=−z2∂ ∂z+ 2jz. (7.52) Similarly, we find that /hatwideJ−=∂ ∂z, (7.53) /hatwideJ3=z∂ ∂z−j. (7.54) These operators obey the su(2) Lie algebra relations [/hatwideJ3,/hatwideJ±] =±/hatwideJ±, [/hatwideJ+,/hatwideJ−] = 2/hatwideJ3, (7.55) and act on the ψm(z) monomials as /hatwideJ3ψm(z) =mψm(z) /hatwideJ±ψm(z) =/radicalbig j(j+ 1)−m(m±1)ψm±1(z). (7.56) This is the familiar action of the su(2) generators on |j,m/angbracketrightbasis states. Exercise 7.1 : Show that with respect to the inner product (7.44) we have /hatwideJ† 3=/hatwideJ3,/hatwideJ† +=/hatwideJ−. 7.2. PHYSICS EXAMPLES 271 Coherent states and the Borel-Weil-Bott theorem We now explain how the spin wavefunctions ψm(z) can be understood as sections of a holomorphic line bundle. Suppose that we have a compact Lie group Gand a unitary irreducible representation g∈G/mapsto→DJ(g). Let|0/angbracketrightbe the normalized highest (or lowest) weight state in the representation space. Consider the stat es |g/angbracketright=DJ(g)|0/angbracketright,/angbracketleftg|=/angbracketleft0|/bracketleftbig DJ(g)/bracketrightbig†. (7.57) The|g/angbracketrightcompose a family of generalized coherent states .2There is a contin- uous infinity of the |g/angbracketright, and so they cannot constitute an orthonormal set on the finite dimensional representation space. The matrix-el ement orthogonal- ity property (6.79), however, provides us us with a useful over-completeness relation I=dim(J) VolG/integraldisplay G|g/angbracketright/angbracketleftg|. (7.58) The integral is over all of G, but many points in Ggive the same contri- bution. The maximal torus Tis the abelian subgroup of Gobtained by exponentiating elements of the Cartan algebra. Because any weight vector is a common eigenvector of the Cartan algebra, elements of Tleave|0/angbracketrightfixed up to a phase. The set of distinct |g/angbracketrightin the integral can therefore be identified withG/T. This coset space is always an even dimensional manifold, an d thus a candidate phase space. Consider in particular the spin- jrepresentation of SU(2). The coset space G/Tis then SU(2) /U(1)/similarequalS2. We can write a general element of SU(2) as U= exp(zJ+) exp(θJ3) exp(γJ−) (7.59) for some complex parameters z,θandγwhich are functions of the three real co-ordinates that parameterize SU(2). We let Uact on the lowest-weight state|j,−j/angbracketright. The rightmost factor has no effect on the lowest weight state , and the middle factor only multiplies it by a constant. We the refore restrict our attention to the states |z/angbracketright= exp(zJ+)|j,−j/angbracketright,/angbracketleftz|=/angbracketleftj,−j|exp(zJ−) = (|z/angbracketright)†. (7.60) 2A. Perelomov, Generalized Coherent States and their Applications , (Springer-Verlag, Berlin 1986). 272 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES These states are not normalized, but have the advantage that the/angbracketleftz|are holomorphic in the parameter z—i.e.they depend on z, but not on z. The set of distinct |z/angbracketrightcan still be identified with the 2-sphere, and z,z are its complex sterographic co-ordinates. This identifica tion is an example of a general property of compact Lie groups: G/T∼=GC/B+. (7.61) HereGCis thecomplexification ofG— the group G, but with its parameters allowed to be complex — and B+is the Borel group whose Lie algebra consists of the Cartan algebra together with the step-up ladder opera tors. The inner product of two |z/angbracketrightstates is /angbracketleftz/prime|z/angbracketright= (1 +zz/prime)2j, (7.62) and the eigenstates |j,m/angbracketrightofJ2andJ3possess coherent state wavefunctions ψ(1) m(z)≡/angbracketleftz|j,m/angbracketright=/radicalBigg 2j! (j−m)!(j+m)!zj+m. (7.63) We recognize these as our spin wavefunctions from the previo us section. The over-completeness relation can be written as I=2j+ 1 2πi/integraldisplaydz∧dz (1 +zz)2j+2|z/angbracketright/angbracketleftz|, (7.64) and provides the inner product for the coherent-state wavef unctions. If ψ(z) =/angbracketleftz|ψ/angbracketrightandχ(z) =/angbracketleftz|χ/angbracketrightthen /angbracketleftψ|χ/angbracketright=2j+ 1 2πi/integraldisplaydz∧dz (1 +zz)2j+2/angbracketleftψ|z/angbracketright/angbracketleftz|χ/angbracketright =2j+ 1 2πi/integraldisplaydz∧dz (1 +zz)2j+2ψ(z)χ(z), (7.65) which coincides with (7.44). The wavefunctions ψ(1) m(z) are singular at the north pole where z=∞. Indeed there is no actual state /angbracketleft∞|because the phase of this putative limiting state would depend on the direction from which we approach th e point at infinity. We may, however, define a second family of coherent s tates |ζ/angbracketright2= exp(ζJ−)|j,j/angbracketright,2/angbracketleftζ|=/angbracketleftj,j|exp(ζJ+), (7.66) 7.2. PHYSICS EXAMPLES 273 and form the wavefunctions ψ(2) m(ζ) =2/angbracketleftζ|j,m/angbracketright. (7.67) These new states and wavefunctions are well defined in the vic inity of the north pole, but singular near the south pole. To find the relation between ψ(2)(ζ) andψ(1)(z) we note that the matrix identity /bracketleftbigg 0−1 1 0/bracketrightbigg/bracketleftbigg 1 0 z1/bracketrightbigg =/bracketleftbigg 1 0 −z−11/bracketrightbigg/bracketleftbigg −z0 0−z−1/bracketrightbigg/bracketleftbigg 1z−1 0 1/bracketrightbigg , (7.68) coupled with the faithfulness of the spin-1 2representation of SU(2), implies the relation ˆwexp(zJ+) = exp (−z−1J−)(−z)2J3exp (z−1J+), (7.69) where ˆw= exp(−iπJ2).We also note that /angbracketleftj,j|ˆw= (−1)2j/angbracketleftj,−j|,/angbracketleftj,−j|ˆw=/angbracketleftj,j|. (7.70) Thus, ψ(1) m(z) =/angbracketleftj,−j|ezJ−|j,m/angbracketright = (−1)2j/angbracketleftj,j|ˆwezJ−|j,m/angbracketright = (−1)2j/angbracketleftj,j|e−z−1J−(−z)2J3ez−1J+|j,m/angbracketright = (−1)2j(−z)2j/angbracketleftj,j|ez−1J+|j,m/angbracketright =z2jψ(2) m(z−1). (7.71) The transition function z2jthat relates ψ(1) m(z) toψ(2) m(ζ≡1/z) depends only onz. We therefore say that the wavefunctions ψ(1) m(z) andψ(2) m(ζ) are the local components of a global section ψm↔|j,m/angbracketrightof aholomorphic line bundle . The requirement that the transition function and its invers e be holomorphic and single valued in the overlap of the zandζcoordinate patches forces 2 j to be an integer. The ψmform a basis for the space of global holomorphic sections of this bundle. Borel, Weil and Bott showed that any finite-dimensional repr esentation of a semi-simple Lie group Gcan be realized as the space of global holomorphic sections of a line bundle over GC/B+. This bundle is constructed from the 274 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES highest (or lowest) weight vectors in the representation by a natural gener- alization of the method we have used for spin. This idea has be en extended by Ed Witten and others to infinite dimensional Lie groups, wh ere it can be used, for example, to quantize two-dimensional gravity. Exercise 7.2 : Normalize the states |z/angbracketright,/angbracketleftz|, by multiplying them by N= (1 +|z|2)−j. Show that N2/angbracketleftz|J3|z/angbracketright=j|z|2−1 |z|2+ 1, N2/angbracketleftz|J+|z/angbracketright=j2z |z|2+ 1, N2/angbracketleftz|J−|z/angbracketright=j2z |z|2+ 1, thus confirming the identification of z,zwith the complex stereographic co- ordinates on the sphere. 7.3 Working in the Total Space We have mostly considered a bundle to be a collection of mathe matical ob- jects attached to a base space, rather than treating the bund le as a geometric object in its own right. In this section we will demonstrate t he advantages to be gained from the latter viewpoint. 7.3.1 Principal Bundles and Associated bundles The fibre bundles that arise in a gauge theory with Lie group Gare called principal G-Bundles , and the fields and wavefunctions are sections of associ- atedbundles. A principal G-bundle comprises the total space, which we here callP, together with the projection, π, to the base space M. The fibre can be regarded as a copy of G π:P→M, π−1(x)∼=G. (7.72) Strictly speaking, the fibre is only required to be a homogene ous space on whichGacts freely and transitively on the right;x→xg. Such a set can be identified with Gafter we have selected a fiducial point f0∈Fto be the group identity. There is no canonical choice for f0and, if the bundle is 7.3. WORKING IN THE TOTAL SPACE 275 twisted, there can be no globally smooth choice. This is beca use a smooth choice forf0in the fibres above an open subset U⊆MmakesPlocally into a product U×G. Being able to extend Uto the entirety of Mmeans thatPis trivial. We will, however, make use of local assignments f0/mapsto→e to introduce bundle co-ordinate charts in which Pis locally a product, and therefore parametrized by ordered pairs ( x,g) withx∈Uandg∈G. To understand the bundles associated withP, it is simplest to define the sections of the associated bundle. Let ϕi(x,g) be a function on the total spacePwith a set of indices icarrying some representation g/mapsto→D(g) of G. We say that ϕi(x,g) is a section of an associated bundle if it varies in a particular way as we run up and down the fibres by acting on them from the rightwith elements of G. We require ϕi(x,gh) =Dij(h−1)ϕj(x,g). (7.73) These sections can be thought of as wavefunctions for a parti cle moving in a gauge field on the base space. The choice of representation Dplays the role of “charge,” and (7.73) are the gauge transformations. Note that we must takeh−1as the argument of Din order for the transformation to be consistent under group multiplication: ϕi(x,gh 1h2) =Dij(h−1 2)ϕj(x,gh 1) =Dij(h−1 2)Djk(h−1 1)ϕk(x,g) =Dik(h−1 2h−1 1)ϕk(x,g) =Dik((h1h1)−1)ϕk(x,g). (7.74) The construction of the associated bundle itself requires r ather more ab- straction. Suppose that the matrices D(g) act on the vector space V. Then the total space PVof the associated bundle consists of equivalence classes ofP×Vunder the relation (( x,g),v)∼((x,gh),D(h−1)v) for all v∈V, (x,g)∈Pandh∈G. The set of G-action equivalence classes in a Cartesian productA×Bis usually denoted by A×GB. Our total space is therefore PV=P×GV. (7.75) We find it conceptually easier to work with the sections as defi ned above, rather than with these equivalence classes. 276 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES 7.3.2 Connections A gauge field is a connection on a principal bundle. The formal definition of a connection is a decomposition of the tangent space TPpofPatp∈Pinto ahorizontal subspace Hp(P) and a vertical subspace Vp(P). We require that Vp(P) be the tangent space to the fibres and Hp(P) to be a complementary subspace, i.e., the direct sum should be the whole tangent space TPp=Hp(P)⊕Vp(P). (7.76) The horizontal subspaces must also be invariant under the pu sh-forward induced from the action on the fibres from the right of a fixed element ofG. More formally, if R[g] :P→Pacts to take p→pg,i.e.by R[g](x,g/prime) = (x,g/primeg), we require R[g]∗Hp(P) =Hpg(P). (7.77) Thus, we get to chose one horizontal subspace in each fibre, th e rest being determined by the right-invariance condition. Given a curve x(t) in the base space we can, by solving the equation ˙g+∂xµ ∂tAµ(x)g= 0, (7.78) liftit to a curve ( x(t),g(t)) in the total space, whose tangent is everywhere horizontal. This lifting operation corresponds to paralle l transporting the initial value g(0) along the curve x(t) to getg(t). TheAµ=iˆλaAa µare a set of Lie-algebra-valued functions that are determined by our ch oice of horizontal subspace. They are defined so that the vector ( δx,−Aµδxµg) is horizontal for each small displacement δxµin the tangent space of M. Here−Aµδxµgis to be understood as the displacement that takes g→(1−Aµδxµ)g. Because we are multiplying Ain from the left, the lifted curve can be slid rigidly up and down the fibres by the right action of any fixed group elem ent. The right-invariance condition is therefore automatically sa tisfied. The directional derivative along the lifted curve is ˙xµDµ= ˙xµ/parenleftBigg/parenleftbigg∂ ∂xµ/parenrightbigg g−Aa µRa/parenrightBigg , (7.79) whereRais a right-invariant vector field on G,i.e., a differential operator on functions defined on the fibres. The Dµare a set of vector fields in TP. These 7.3. WORKING IN THE TOTAL SPACE 277 covariant derivatives span the horizontal subspace at each point p∈P, and have Lie brackets [Dµ,Dν] =−Fa µνRa. (7.80) HereFµν, is given in terms of the structure constants appearing in th e Lie brackets [Ra,Rb] =fc abRcby Fc µν=∂µAc ν−∂νAc µ−fc abAa µAb ν. (7.81) We can also write Fµν=∂µAν−∂νAµ+ [Aµ,Aν]. (7.82) whereFµν=iˆλaFa µνand [ˆλa,ˆλb] =ifc abˆλc. Because the Lie bracket of the Dµis a linear combination of the Ra, it lies entirely in the vertical subspace. Consequently, when Fµν/negationslash= 0, theDµare not in involution, and Frobenius’ theorem tells us that the hori zontal subspaces cannot fit together to form the tangent spaces to a smooth foli ation ofP. We make contact with the more familiar definitions of covaria nt deriva- tives by remembering that right invariant vector fields are derivatives that involve infinitesimal multiplication from the left. Their definition is Raϕi(x,g) = lim /epsilon1→01 /epsilon1/parenleftBig ϕi(x,(1 +i/epsilon1ˆλa)g)−ϕi(x,g)/parenrightBig , (7.83) where [ ˆλa,ˆλb] =ifc abˆλc. Sinceϕi(x,g) is a section of the associated bundle, we know how it varies when we multiply group elements in on the right. We therefore write (1 +i/epsilon1ˆλa)g=gg−1(1 +i/epsilon1ˆλa)g, (7.84) and from this, (and writing gforD(g) where it makes for compact notation) we find Raϕi(x,g) = lim /epsilon1→0/parenleftBig Dij(g−1(1−i/epsilon1ˆλa)g)ϕj(x,g)−ϕi(x,g)/parenrightBig //epsilon1 =−Dij(g−1)(iˆλa)jkDkl(g)ϕl(x,g) =−i(g−1ˆλag)ijϕj. (7.85) Herei(ˆλa)ijis the matrix representing the Lie algebra generator iˆλain the representation g/mapsto→D(g). Acting on sections, we therefore have Dµϕ= (∂µϕ)g+ (g−1Aµg)ϕ. (7.86) 278 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES This still does not look too familiar because the derivative s with respect to xµare being taken at fixedg. We normally fix a gauge by making a choice of g=σ(x) for eachxµ. The conventional wavefunction ϕ(x) is thenϕ(x,σ(x)). We can use ϕ(x,σ(x)) =σ−1(x)ϕ(x,e), to obtain ∂µϕ= (∂µϕ)σ+/parenleftbig ∂µσ−1/parenrightbig σϕ= (∂µϕ)σ−/parenleftbig σ−1∂µσ/parenrightbig ϕ. (7.87) From this we get a derivative ∇µdef=∂µ+ (σ−1Aµσ+σ−1∂µσ) =∂µ+Aµ. (7.88) on functions ϕ(x)≡ϕ(x,σ(x)) defined (locally) on the base space M. This is the conventional covariant derivative, now containing g auge fields Aµ(x) that are gauge transformations of our g-independentAµ. The derivative has been constructed so that ∇µϕ(x) =Dµϕ(x,g)|g=σ(x), (7.89) and has commutator [∇µ,∇ν] =σ−1Fµνσ=Fµν. (7.90) Note the sign change vis-a-vis equation (7.80). It is the curvature tensor Fµνthat we have met previously. Recall that it provides a Lie algebra valued two-form F≡1 2Fµνdxµdxν=dA+A2(7.91) on the base space. The connection A≡Aµdxµis a one-form on the base space, and both FandAhave been defined only in the region U⊂Mwhere the smooth gauge-choice section σ(x) has been selected. 7.3.3 Monople harmonics The total-space operations and definitions seem rather abst ract. We demon- strate their power by solving the Schr¨ odinger problem for a charged particle confined to a unit sphere surrounding a magnetic monopole. Th e conven- tional approach to this problem involves first selecting a ga uge for vector the potential A, which, because of the monopole, is necessarily singular at a 7.3. WORKING IN THE TOTAL SPACE 279 Dirac string located somewhere on the sphere, and then delvi ng into prop- erties of Gegenbauer polynomials. Eventually we find the gau ge-dependent wavefunction. By working with the total space, however, we c an solve the problem in all gauges at once , and the problem becomes a simple exercise in Lie group geometry. Recall that the SU(2) representation matrices DJ mn(θ,φ,ψ ) form a com- plete orthonormal set of functions on the group manifold S3. There will be a similar complete orthonormal set of representation matric es on the manifold of any compact Lie group G. Given a subgroup H∈G, we will use these matrices to construct bundles associated to a principal H-bundle that has G as its total space, and the coset space G/H as its base space. The fibres will be copies of H, and the projection πthe usual projection G→G/H. The functions DJ(g) are not in general functions on the coset space G/H as they depend on the choice of representative. Instead, bec ause of the representation property, they vary with the choice of re presentative in a well-defined way, DJ mn(gh) =DJ mn/prime(g)DJ n/primen(h). (7.92) Since we are dealing with compact groups, the representatio ns can be taken to be unitary and [DJ mn(gh)]∗= [DJ mn/prime(g)]∗[DJ n/primen(h)]∗(7.93) =DJ nn/prime(h−1)[DJ mn/prime(g)]∗. (7.94) This is the correct variation under the right action of the gr oupHfor the set of functions [ DJ mn(gh)]∗to be sections of a bundle associated with the principal fibre bundle G→G/H. The representation h/mapsto→D(h) ofHis not necessarily that defined by the label Jbecause irreducible representations of Gmay be reducible under H;Ddepends on what representation of Hthe indexnbelongs to. If Dis the identity representation, then the functions are functions on G/Hin the ordinary sense. For G= SU(2) and HtheU(1) subgroup generated by J3, the quotient space is just S2, and projection is the Hopf map: S3→S2. The resulting bundle can be called the Hopf bundle. It is not a really new object however, because it is a generali zation of the monopole bundle of the preceding section. Parameterizing S U(2) with Euler angles, so that DJ mn(θ,φ,ψ ) =/angbracketleftJ,m|e−iφJ3e−iθJ2e−iψJ3|J,n/angbracketright, (7.95) 280 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES shows that the Hopf map consists of simply forgetting about ψ, so Hopf : [(θ,φ,ψ )∈S3]/mapsto→[(θ,φ)∈S2]. (7.96) The bundle is twisted because S3is not a product S2×S1. Takingn= 0 gives us functions independent of ψ, and we obtain the well-known identification of the spherical harmonics with representation matrices YL m(θ,φ) =/radicalbigg 2L+ 1 4π[D(L) m0(θ,φ,0)]∗. (7.97) Forn= Λ/negationslash= 0 we get sections of a bundle with Chern number 2Λ. These sections are the monopole harmonics YJ m;Λ(θ,φ,ψ ) =/radicalbigg 2J+ 1 4π[DJ mΛ(θ,φ,ψ )]∗(7.98) for a monopole of flux/integraltext eBd(Area) = 4 πΛ. The integrality of the Chern number tells us that the flux 4 πΛ must be an integer multiple of 2 π. This gives us a geometric reason for why the eigenvalues mofJ3can only be an integer or half integer. The monopole harmonics have a non-trivial ∝eiψΛdependence on the choice we make for ψat each point on S2, and we cannot make a globally smooth choice; we always encounter a point where there is a si ngularity. These sections of the twisted bundle have to be constructed i n patches and glued together transition functions. We now show that the monopole harmonics are eigenfunctions o f the Schr¨ odinger operator, −∇2, containing the gauge field connection, just as the spherical harmonics are eigenfunctions of the Laplacian on the sphere. This is a simple geometrical exercise. Because they are irreduci ble representations, theDJ(g) are automatically eigenfunctions of the quadratic Casimi r operator (J2 1+J2 2+J2 3)DJ(g) =J(J+ 1)DJ(g). (7.99) TheJican be either right or left-invariant vector fields on G; the quadratic Casimir is the same second-order differential operator in ei ther case, and it is a good guess that it is proportional to the Laplacian on the group mani- fold. Taking a locally geodesic co-ordinate system (in whic h the connection vanishes) confirms this: J2=−∇2on the three-sphere. The operator in (7.99) is not the Laplacian we want, however. What we need is t he∇2on 7.3. WORKING IN THE TOTAL SPACE 281 the two-sphere S2=G/H, including the the connection. This ∇2operator differs from the one on the total space since it must contain on ly differential operators lying in the horizontal subspaces. There is a natu ral notion of or- thogonality in the Lie group, deriving from the Killing form , and it is natural to choose the horizontal subspaces to be orthogonal to the fib res ofG/H. Since multiplication on the right by the subgroup generated byJ3moves one up and down the fibres, the orthogonal displacements are o btained by multiplication on the right by infinitesimal elements made b y exponentiating J1andJ2. The desired∇2is thus made out of the left-invariant vector fields (which act by multiplication on the right), J1andJ2only. The wave operator must be −∇2=J2 1+J2 2=J2−J2 3. (7.100) Applying this to the YJ m;Λwe see that they are eigenfunctions of −∇2onS2 with eigenvalues J(J+ 1)−Λ2. The Laplace eigenvalues for our flux 4 πΛ monopole problem are therefore EJ,m= (J(J+ 1)−Λ2), J≥|Λ|,−J≤m≤J. (7.101) The utility of the monopole Harmonics is not restricted to ex otic monopole physics. They occur in molecular and nuclear physics as the w avefunctions for the rotational degrees of freedom of diatomic molecules and uniaxially deformed nuclei that possess angular momentum Λ about their axis of sym- metry.3 Exercise 7.3 : Compare these energy levels for a particle on a sphere with t hose of the Landau level problem on the plane. Show that for any fixe d flux the low-lying energies remain close to E= (eB/m particle )(n+ 1/2),nzero or a positive integer, but their degeneracy is is equal to the num ber of flux units penetrating the sphere plus one . 7.3.4 Bundle connection and curvature forms Recall that in section 7.3.2 we introduced the Lie-Algebra- valued functions Aµ(x). We now use these functions to introduce the bundle connection form Athat lives in T∗P. We set A=Aµdxµ(7.102) 3This is explained, with chararacteristic terseness, in a fo otnote on page 317 of Landau and Lifshitz’ Quantum Mechanics (Third Edition). 282 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES and Adef=g−1/parenleftbig A+δgg−1/parenrightbig g. (7.103) In these definitions, xandgare the local co-ordinates in which points in the total space are labelled as ( x,g), anddacts on functions of x, and the “δ” is used to denote the exterior derivative acting on the fibre .4We have, then, that δxµ= 0 anddg= 0. The combinations δgg−1andg−1δgare respectively the right- and left-invariant Maurer-Cartan form on the group. The complete exterior derivative in the total space require s us to differen- tiate both with respect to gand with respect to x, and is given by dtot=d+δ. Becaused2,δ2and (d+δ)2=d2+δ2+dδ+δdare all zero, we must have δd+dδ= 0. (7.104) We now define the bundle curvature form in terms of Ato be Fdef=dtotA+A2. (7.105) To compute Fin terms ofA(x) andgwe need the ingredients dA=g−1(dA)g, (7.106) and δA=−(g−1δg)A−A(g−1δg)−(g−1δg)2. (7.107) We find that F= (d+δ)A+A2=g−1/parenleftbig dA+A2/parenrightbig g =g−1Fg, (7.108) where F=1 2Fµνdxµdxν, (7.109) and Fµν=∂µAν−∂νAµ+ [Aµ,Aν]. (7.110) Although we have defined the connection form Ain terms of the local bundle co-ordinates ( x,g), it is, in fact, an intrinsic quantity, i.e.it is has a global existence independent of the choice of these co-ordi nates. Ahas been constructed so that 4It isnottherefore to be confused with the Hodge δ=d†operator. 7.3. WORKING IN THE TOTAL SPACE 283 •A vector is annihilated by Aif and only if it is horizontal. In particular A(Dµ) = 0 for all covariant derivatives Dµ. •The connection form is constant on left-invariant vector fields on the fibres. In particular A(La) =iˆλa. Between them, the globally defined fields Dµ∈Hp(P) andLa∈Vp(P) span the tangent space TPp. Consequently the two properties listed above tell us how to evaluate Aon any vector, and so define it uniquely and globally. From the globally defined and gauge invariant Aand its associated cur- vature F, and for any local gauge-choice section σ: (U⊂M)→P, we can recover the gauge-dependent base-space forms AandFas the pull-backs A=σ∗A, F =σ∗F, (7.111) toU⊂Mof the total-space forms. The resulting forms are A=/parenleftbig σ−1Aµσ+σ−1∂µσ/parenrightbig dxµ, F =1 2/parenleftbig σ−1Fµνσ/parenrightbig dxµdxν, (7.112) and coincide with the equations connecting AµwithAµandFµνwithFµν that we obtained in section 7.3.2. We should take care to note that thedxµ that appear in AandFare differential forms on M, while the dxµthat appear inAandFare differential forms on P. Now the projection πis a left inverse of the gauge-choice section σ,i.e.π◦σ= identity. The associated pull-backs are also inverses, but with the order reversed: σ∗◦π∗= identity. These maps relate the two sets of “ dxµ” by dxµ|M=σ∗(dxµ|P),ordxµ|P=π∗(dxµ|M). (7.113) We now explain the advantage of knowing the total space conne ction and curvature forms. Consider the Chern character ∝trF2on the base-space M. We can use the bundle projection πto pull this form back to total space. From Fµν= (gσ−1)−1Fµν(gσ−1), (7.114) we find that π∗/parenleftbig trF2/parenrightbig = trF2. (7.115) Now A,Fanddtothave the same calculus properties as A,Fandd. The manipulations that give trF2=dtr/parenleftbigg AdA+2 3A3/parenrightbigg 284 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES also show, therefore, that trF2=dtottr/parenleftbigg AdtotA+2 3A3/parenrightbigg . (7.116) There is a big difference in the significance of the computatio n, however. The bundle connection Ais globally defined. Consequently, the form ω3(A)≡tr/parenleftbigg AdtotA+2 3A3/parenrightbigg (7.117) is also globally defined. The pull-back to the total space of t he Chern char- acter isdtotexact! This miracle works for all characteristic classes: o n the base-space they are exact only when the bundle is trivial; on the total space they are always exact. We have seen this phonomenon before, for example in exercise 6.7. The area formd[Area] = sin θdθdφ is closed but not exact on S2. When pulled back toS3by the Hopf map, the area form becomes exact: Hopf∗d[Area] = sin θdθdφ =d(−cosθdφ+dψ). (7.118) 7.3.5 Characteristic classes as obstructions The generalized Gauss-Bonnet theorem states that, for a com pact orientable even-dimensional manifold M, the integral of the Euler class over Mis equal to the Euler character χ(M). Shiing-Shen Chern used the exactness of the pull-back of the Euler class to give an elegant intrinsic pro of5of this theorem. Chern showed that the integral of the Euler class over Mwas equal to the sum of the Poincare-Hopf indices of any tangent vector field o nM, a sum we independently know to equal the Euler character χ(M). We illustrate his strategy by showing how a non-zero ch 2(F) provides a similar index sum for the singularities of any section of an SU(2)-bundle over a fo ur-dimensional base space. This result provides an interpretation of chara cteristic classes as obstructions to the existence of global sections. Letσ:M→Pbe a section of an SU(2) principal bundle Pover a four-dimensional compact orientable manifold Mwithout boundary. For any SU(n) group we have ch 1(F)≡0, but /integraldisplay Mch2(F) =−1 8π2/integraldisplay Mtr(F2) =n, (7.119) 5S-J. Chern, Ann. Math. 47(1946) 85-121. This paper is a readable classic. 7.3. WORKING IN THE TOTAL SPACE 285 can be non-zero. The section σwill, in general, have points xiwhere it becomes singular. We punch infinitesimal holes in Msurrounding the singular points. The manifoldM/prime= (M\holes) will have as its boundary ∂M/primea disjoint union of small three-spheres. We denote by Σ the image of M/primeunder the map σ:M/prime→P. This Σ will be a submanifold of P, whose boundary will be equal in homology to a linear combination of the boundary com ponents of M/primewith integer coefficients. We show that the Chern number nis equal to the sum of these coefficients. We begin by using the projection πto pull back ch 2(F), to the bundle, where we know that π∗ch2(F) =−1 8π2dtotω3(A). (7.120) Now we can decompose ω3(A) into terms of different bi-degree, i.e.into terms that are p-forms indandq-forms inδ. ω3(A) =ω0 3+ω1 2+ω2 1+ω3 0. (7.121) Here the superscript counts the form-degree in δ, and the subscript the form- degree ind. The only term we need to know explicitly is ω3 0. This comes from theg−1δgpart of A, and is ω3 0= tr/parenleftbigg (g−1δg)δ(g−1δg) +2 3(g−1δg)3/parenrightbigg = tr/parenleftbigg −(g−1δg)3+2 3(g−1δg)3/parenrightbigg =−1 3(g−1δg)3. (7.122) We next use the map σ:M/prime→Pto pull the right-hand side of (7.120) back fromPtoM/prime. We recall that acting on forms on M/primewe haveσ∗◦π∗= identity. Thus /integraldisplay Mch2(F) =/integraldisplay M/primech2(F) =/integraldisplay M/primeσ∗◦π∗ch2(F) =−1 8π2/integraldisplay M/primeσ∗dtotω3(A) =−1 8π2/integraldisplay Σdtotω3(A) 286 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES =−1 8π2/integraldisplay ∂Σω3(A) =1 24π2/integraldisplay ∂Σ(g−1δg)3. (7.123) At the first step we have observed that the omitted spheres mak e a negligeable contribution to the integral over M, and at the last step we have used the fact that the boundary of Σ, has significant extent only along the fibres, so all contributions to the integral over ∂Σ come from the purely vertical component of ω3(A), which isω3 0=−1 3(g−1dg). We know (see exercise 6.8) that for maps g/mapsto→U∈SU(2) we have /integraldisplay tr (g−1dg)3= 24π2×winding number We conclude that /integraldisplay Mch2(F) =1 24π2/integraldisplay ∂Σ(g−1δg)3=/summationdisplay singularities xiNi (7.124) whereNiis the Brouwer degree of the map σ:S3→SU(2)∼=S3on the small sphere surrounding xi. It turns out that for any SU( n) the integral of tr( g−1δg)3is 24π2times an integer winding number of gabout homology spheres. The second Chern number of a SU( n)-bundle is therefore also equal to the sum of the winding- number indices of the section about its singularities. Cher n’s strategy can be used to relate other characteristic classes to obstructi ons to the existence of global sections of appropriate bundles. 7.3.6 Stora-Zumino descent equations In the previous sections we met the forms A=g−1Ag+g−1δg (7.125) and A=σ−1Aσ+σ−1dσ. (7.126) The group element glabeled points on the fibres and was independent x, whileσ(x) was the gauge-choice section of the bundle and depended on x. 7.3. WORKING IN THE TOTAL SPACE 287 The two quantities AandAlook similar, but are not identical. A third superficially similar but distinct object is met with in the B RST (Becchi- Rouet-Stora-Tyutin) approach to quantizing gauge theorie s, and also in the geometric theory of anomalies. We describe it here to alert t he reader to the potential for confusion. Rather than attempting to define this new differential form ri gorously, we will first explain how to calculate with it, and only then in dicate what it is. We begin by considering a fixed connection form AonM, and its orbit under the action of the group Gof gauge transformations. This elements of this infinite dimensional group are maps g:M→Gequipped with pointwise productg1g2(x) =g1(x)g2(x). Thisg(x) is neither the fibre co-ordinate g, nor the gauge choice section σ(x). The gauge transformation g(x) acts onA to giveAgwhere Ag=g−1Ag+g−1dg. (7.127) We now introduce an object v(x) =g−1δg, (7.128) and consider A=Ag+v=g−1Ag+g−1dg+g−1δg. (7.129) This 1-form appears to be a hybrid of the earlier quantities, but we will see that it has to be considered as something new. The essenti al difference from what has gone before is that we want vto behave like g−1δg, in that δv=−v2, and yet to depend on x. In particular we want δto behave as an exterior derivative that implements an infinitesimal gau ge transformation that takesg→g+δg. Thus, δ(g−1dg) =−(g−1δg)(g−1dg) +g−1δdg =−(g−1δg)(g−1dg)−(g−1dg)(g−1δg) + (g−1dg)(g−1δg)−g−1dδg =−v(g−1dg)−(g−1dg)v−dv, (7.130) and hence δAg=−vAg−Agv−dv. (7.131) Previously g−1dg≡0, and so there was no “ dv” inδ(gauge field). We can define a curvature associated with A Fdef=dtotA+A2, (7.132) 288 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES and compute F= (d+δ)(Ag+v) + (Ag+v)2 =dAg+dv+δAg+δv+ (Ag)2+Agv+vAg+v2 =dAg+ (Ag)2 =g−1Fg, (7.133) Stora calls (7.133) the Russian formula . Because Fis yet another gauge transform of F, we have trF2= trF2= (d+δ) tr/parenleftbigg A(d+δ)A+2 3A3/parenrightbigg (7.134) and can decompose the right-hand side into terms that are sim ultaneously p-foms indandq-forms inδ. The left hand side, tr F2= trF2, of (7.134) is independent of v. The right hand side of (7.134) contains ω3(A) which we expand as ω3(Ag+v) =ω0 3(Ag) +ω1 2(v,Ag) +ω2 1(v,Ag) +ω3 0(v). (7.135) As in the previous section, the superscript counts the form- degree inδ, and the subscript the form-degree in d. Explicit computation shows that ω0 3(Ag) = tr/parenleftbig AgdAg+2 3(Ag)3/parenrightbig , ω1 2(v,Ag) = tr ( vdAg), ω2 1(v,Ag) =−tr (Agv2), ω3 0(v) =−1 3v3(7.136) For example, ω3 0(v) = tr/parenleftbigg vδv+2 3v3/parenrightbigg = tr/parenleftbigg v(−v2) +2 3v3/parenrightbigg =−1 3v3. (7.137) With this decomposition, (7.116) falls apart into the chain ofdescent equa- tions trF2=dω0 3(Ag), δω0 3(Ag) =−dω1 2(v,Ag), δω1 2(v,Ag) =−dω2 1(v,Ag), δω2 1(v,Ag) =−dω3 0(v), δω3 0(v) = 0. (7.138) 7.3. WORKING IN THE TOTAL SPACE 289 Let us verify, for example, the penultimate equation δω2 1(v,Ag) =−dω3 0(v). The left-hand side is −δtr (Agv2) =−tr (−Av3−vAgv2−dvv2) = tr (dvv2), (7.139) the terms involving Aghaving cancelled viathe cyclic property of the trace and the fact that Aganticommutes with v. The right-hand side is −d/parenleftbig −1 3trv3/parenrightbig = tr (dvv2) (7.140) as required. The descent equations were introduced by Raymond Stora and B runo Zu- mino as a tool for obtaining and systematizing information a boutanomalies in the quantum field theory of fermions interacting with the g auge fieldAg. Theωq p(v,Ag) arep-forms in the dxµ, and before use they are integrated over p-cycles inM. This process is understood to produce local functionals of Ag that remain q-forms inδg. For example, in 2 nspace-time dimensions, the integral I[g−1δg,Ag] =/integraldisplay Mω1 2n(g−1δg,Ag) (7.141) has the properties required for it to be a candidate for the an omalous vari- ationδS[Ag] of the fermion effective action due to an infinitesimal gauge transformation g→g+δg. In particular, when ∂M=∅, we have δI[g−1δg,Ag] =/integraldisplay Mδω1 2n(v,Ag) =−/integraldisplay Mdω2 2n−1(v,Ag) = 0. (7.142) This is the Wess-Zumino consistency condition thatδ(δS) must obey as a consequence of δ2= 0. In addition to producing a convenient solution of the Wess-Z umino condi- tion, the descent equations provide a compact derivation of the gauge trans- formation properties of useful differential forms. We will n ot seek to explain further the physical meaning of these forms, leaving this to a field theory course. The similarity between AandAlead various authors to attempt to iden- tify them, and in particular to identify v(x) with the g−1δgMaurer-cartan form appearing in A. However the physical meaning of expressions such as d(g−1δg) precludes such a simple interpretation. In evaluating dv∼d(g−1δg) on a vector field ξa(x)Larepresenting an infinitesimal gauge transformation, 290 CHAPTER 7. THE GEOMETRY OF FIBRE BUNDLES we are to first to insert the field into v∼g−1δgto obtain the xdependent Lie algebra element iξa(x)ˆλa, and only then to take the exterior derivative to obtainiˆλa∂µξadxµ. The result therefore involves derivatives of the com- ponentsξa(x). The evaluation of an ordinary differential form on a vector field never produces derivatives of the vector components. To understand what the Stora-Zumino forms are, imagine that we equip a two dimensional fibre bundle E=M×Fwith base-space co-ordinate xand fibre co-ordinate y. Ap= 1,q= 1 form on Ewill then be F=f(x,y)dxδy for some function f(x,y). There is only one object δy, and there is no meaning to integrating Foverxto leave a 1-form in δyonE. The space of forms introduced by Stora and Zumino, on the other hand, wo uld contain elements such as J=/integraldisplay Mj(x,y)dxδyx (7.143) where there is a distinct δyxfor eachx∈M. If we take, for example, j(x,y) =δ/prime(x−a). we evaluate Jon the vector field Y(x,y)∂yas J[Y(x,y)∂y] =/integraldisplay δ/prime(x−a)Y(x,y)dx=−Y/prime(a,y). (7.144) The conclusion is that that the 1-form form field v(x)∼g−1δgmust be considered as the left-invariant Maurer-Cartan form on the infinite dimen- sional Lie groupG, rather than a Maurer-Cartan form on the finite dimen- sional Lie group G. The/integraltext Mωq 2n(v,Ag) are therefore elements of the coho- mology group Hq(AG) of theGorbit ofA, a rather complicated object. For a thorough discussion see: J. A. de Azc´ arraga, J. M. Izquier do,Lie groups, Lie Algebras, Cohomology and some Applications in Physics , published by Cambridge University Press. Chapter 8 Complex Analysis I Although this chapter is called complex analysis , we will try to develop the subject as complex calculus — meaning that we shall follow the calculus course tradition of telling you how to do things, and explain ing why theorems are true, with arguments that would not pass for rigorous pro ofs in a course on real analysis. We try, however, to tell no lies. This chapter will focus on the basic ideas that need to be unde rstood before we apply complex methods to evaluating integrals, an alysing data, and solving differential equations. 8.1 Cauchy-Riemann equations We focus on functions, f(z), of a single complex variable, z, wherez=x+iy. We can think of these as being complex valued functions of two real variables, xandy. For example f(z) = sinz≡sin(x+iy) = sinxcosiy+ cosxsiniy = sinxcoshy+icosxsinhy. (8.1) Here, we have used sinx=1 2i/parenleftbig eix−e−ix/parenrightbig ,sinhx=1 2/parenleftbig ex−e−x/parenrightbig , cosx=1 2/parenleftbig eix+e−ix/parenrightbig ,coshx=1 2/parenleftbig ex+e−x/parenrightbig , 291 292 CHAPTER 8. COMPLEX ANALYSIS I to make the connection between the circular and hyperbolic f unctions. We shall often write f(z) =u+iv, whereuandvare real functions of xandy. In the present example, u= sinxcoshyandv= cosxsinhy. If all four partial derivatives ∂u ∂x,∂v ∂y,∂v ∂x,∂u ∂y, (8.2) exist and are continuous then f=u+ivis differentiable as a complex- valued function of two real variables. This means that we can approximate the variation in fas δf=∂f ∂xδx+∂f ∂yδy+···, (8.3) where the dots represent a remainder that goes to zero faster than linearly asδx,δygo to zero. We now regroup the terms, setting δz=δx+iδy, δz=δx−iδy, so that δf=∂f ∂zδz+∂f ∂zδz+···, (8.4) where we have defined ∂f ∂z≡1 2/parenleftbigg∂f ∂x−i∂f ∂y/parenrightbigg , ∂f ∂z≡1 2/parenleftbigg∂f ∂x+i∂f ∂y/parenrightbigg . (8.5) Now our function f(z) does not depend on z, and so it must satisfy ∂f ∂z= 0. (8.6) Thus, with f=u+iv, 1 2/parenleftbigg∂ ∂x+i∂ ∂y/parenrightbigg (u+iv) = 0 (8.7) i.e. /parenleftbigg∂u ∂x−∂v ∂y/parenrightbigg +i/parenleftbigg∂v ∂x+∂u ∂y/parenrightbigg = 0. (8.8) 8.1. CAUCHY-RIEMANN EQUATIONS 293 Since the vanishing of a complex number requires the real and imaginary parts to be separately zero, this implies that ∂u ∂x= +∂v ∂y, ∂v ∂x=−∂u ∂y. (8.9) These two relations between uandvare known as the Cauchy-Riemann equations , although they were probably discovered by Gauss. If our con tinu- ous partial derivatives satisfy the Cauchy-Riemann equati ons atz0=x0+iy0 then we say that the function is complex differentiable (or just differentiable) at that point. By taking δz=z−z0, we have δf≡f(z)−f(z0) =∂f ∂z(z−z0) +···, (8.10) where the remainder, represented by the dots, tends to zero f aster than|z−z0| asz→z0. This validity of this linear approximation to the variatio n inf(z) is equivalent to the statement that the ratio f(z)−f(z0) z−z0(8.11) tends to a definite limit as z→z0from any direction. It is the direction- independence of this limit that provides a proper meaning to the phrase “does not depend on z.” Since we are not allowing dependence on ¯ z, it is natural to drop the partial derivative signs and write the li mit as an ordinary derivative lim z→z0f(z)−f(z0) z−z0=df dz. (8.12) We will also use Newton’s fluxion notation df dz≡f/prime(z). (8.13) The complex derivative obeys exactly the same calculus rule s as ordinary real derivatives: d dzzn=nzn−1, d dzsinz= cosz, d dz(fg) =df dzg+fdg dz, etc. (8.14) 294 CHAPTER 8. COMPLEX ANALYSIS I If the function is differentiable at all points in an arcwise- connected1open set, or domain ,D, the function is said to be analytic there. The words regular orholomorphic are also used. 8.1.1 Conjugate pairs The functions uandvcomprising the real and imaginary parts of an analytic function are said to form a pair of harmonic conjugate functions . Such pairs have many properties that are useful for solving physical pr oblems. From the Cauchy-Riemann equations we deduce that/parenleftbigg∂2 ∂x2+∂2 ∂y2/parenrightbigg u= 0, /parenleftbigg∂2 ∂x2+∂2 ∂y2/parenrightbigg v= 0. (8.15) and so both the real and imaginary parts of f(z) are automatically harmonic functions of x,y. Further, from the Cauchy-Riemann conditions, we deduce tha t ∂u ∂x∂v ∂x+∂u ∂y∂v ∂y= 0. (8.16) This means that ∇u·∇v= 0. We conclude that, provided that neither of these gradients vanishes, the pair of curves u=const. andv=const. intersect at right angles. If we regard uas the potential φsolving some electrostatics problem ∇2φ= 0, then the curves v=const. are the associated field lines. Another application is to fluid mechanics. If vis the velocity field of an irrotational (∇×v=0) flow, then we can (perhaps only locally) write the flow field as a gradient vx=∂xφ, vy=∂yφ, (8.17) whereφis avelocity potential . If the flow is incompressible ( ∇·v= 0), then we can (locally) write it as a curl vx=∂yχ, vy=−∂xχ, (8.18) 1Arcwise connected means that any two points in Dcan be joined by a continuous path that lies wholely within D. 8.1. CAUCHY-RIEMANN EQUATIONS 295 whereχis astream function . The curves χ=const. are the flow streamlines. If the flow is both irrotational and incompressible, then we m ay use either φ orχto represent the flow, and, since the two representations mus t agree, we have ∂xφ= +∂yχ, ∂yφ=−∂xχ. (8.19) Thusφandχare harmonic conjugates, and so the complex combination Φ =φ+iχis an analytic function called the complex stream function . A conjugate vexists (at least locally) for any harmonic function u. To see why, assume first that we have a ( u,v) pair obeying the Cauchy-Riemann equations. Then we can write dv=∂v ∂xdx+∂v ∂ydy =−∂u ∂ydx+∂u ∂xdy. (8.20) This observation suggests that if we are given a harmonic fun ctionuin some simply connected domain D, we can define avby setting v(z) =/integraldisplayz z0/parenleftbigg −∂u ∂ydx+∂u ∂xdy/parenrightbigg +v(z0), (8.21) for some real constant v(z0) and point z0. The integral does not depend on choice of path from z0toz, and sov(z) is well defined. The path indepen- dence comes about because the curl ∂ ∂y/parenleftbigg −∂u ∂y/parenrightbigg −∂ ∂x/parenleftbigg∂u ∂x/parenrightbigg =−∇2u (8.22) vanishes, and because in a simply connected domain all paths connecting the same endpoints are homologous. We now verify that this candidate v(z) satisfies the Cauchy-Riemann realtions. The path independence, allows us to make our final approach to z=x+iyalong a straight line segment lying on either the xoryaxis. If we approach along the xaxis, we have v(z) =/integraldisplayx/parenleftbigg −∂u ∂y/parenrightbigg dx/prime+ rest of integral, (8.23) 296 CHAPTER 8. COMPLEX ANALYSIS I and may use d dx/integraldisplayx f(x/prime,y)dx/prime=f(x,y) (8.24) to see that∂v ∂x=−∂u ∂y(8.25) at (x,y). If, instead, we approach along the yaxis, we may similarly compute ∂v ∂y=∂u ∂x. (8.26) Thusv(z) does indeed obey the Cauchy-Riemann equations. Because of the utility the harmonic conjugate it is worth giv ing a practical recipe for finding it, and so obtaining f(z) when given only its real part u(x,y). The method we give below is one we learned from John d’Angel o. It is more efficient than those given in most textbooks. We first observe that iffis a function of zonly, thenf(z) depends only on z. We can therefore define a function fofzby settingf(z) =f(z). Now 1 2/parenleftBig f(z) +f(z)/parenrightBig =u(x,y). (8.27) Set x=1 2(z+z), y=1 2i(z−z), (8.28) so u/parenleftbigg1 2(z+z),1 2i(z−z)/parenrightbigg =1 2/parenleftbig f(z) +f(z)/parenrightbig . (8.29) Now setz= 0, while keeping zfixed! Thus f(z) +f(0) = 2u/parenleftBigz 2,z 2i/parenrightBig . (8.30) The function fis not completely determined of course, because we can alway s add a constant to v, and so we have the result f(z) = 2u/parenleftBigz 2,z 2i/parenrightBig +iC, C∈R. (8.31) For example, let u=x2−y2. We find f(z) +f(0) = 2/parenleftBigz 2/parenrightBig2 −2/parenleftBigz 2i/parenrightBig2 =z2, (8.32) 8.1. CAUCHY-RIEMANN EQUATIONS 297 or f(z) =z2+iC, C∈R. (8.33) The business of setting setting z= 0, while keeping zfixed, may feel like a dirty trick, but it can be justified by the (as yet to be proved ) fact that f has a convergent expansion as a power series in z=x+iy. In this expansion it is meaningful to let xandythemselves be complex, and so allow zand zto become two independent complex variables. Anyway, you ca n always check ex post facto that your answer is correct. 8.1.2 Conformal Mapping An analytic function w=f(z) maps subsets of its domain of definition in the “z” plane on to subsets in the “ w” plane. These maps are often useful for solving problems in two dimensional electrostatics or fl uid flow. Their simplest property is geometrical: such maps are conformal . Z Z1 0 1−Z Z1Z 1−Z1 1−ZZ−1Z Figure 8.1: An illustration of conformal mapping. The unshaded “triang le” markedzis mapped into the other five unshaded regions by the function s labeling them. Observe that although the regions are distor ted, the angles of the “triangle” are preserved by the maps (with the exception of those corners that get mapped to infinity). Suppose that the derivative of f(z) at a point z0is non-zero. Then, for z nearz0we have f(z)−f(z0)≈A(z−z0), (8.34) 298 CHAPTER 8. COMPLEX ANALYSIS I where A=df dz/vextendsingle/vextendsingle/vextendsingle/vextendsingle z0. (8.35) If you think about the geometric interpretation of complex m ultiplication (multiply the magnitudes, add the arguments) you will see th at the “f” image of a small neighbourhood of z0is stretched by a factor |A|, and rotated through an angle arg A— but relative angles are not altered. The map z/mapsto→ f(z) =wis therefore isogonal . Our map also preserves orientation (the sense of rotation of the relative angle) and these two properties, isogonality and orientation-preservation, are what make the map conformal .2The conformal property fails at points where the derivative vanishes or be comes infinite. If we can find a conformal map z(≡x+iy)/mapsto→w(≡u+iv) of some domainDto another D/primethen a function f(z) that solves a potential theory problem (a Dirichlet boundary-value problem, for example) inDwill lead to f(z(w)) solving an analogous problem in D/prime. Consider, for example, the map z/mapsto→w=z+ez. This map takes the strip−∞<x<∞,−π≤y≤πto the entire complex plane with cuts from −∞+iπto−1 +iπand from−∞−iπto−1−iπ. The cuts occur because the images of the lines y=±πget folded back on themselves at w=−1±iπ, where the derivative of w(z) vanishes. (See figure 8.2) In this case, the imaginary part of the function f(z) =x+iytrivially solves the Dirichlet problem ∇2 x,yy= 0 in the infinite strip, with y=π on the upper boundary and y=−πon the lower boundary. The function y(u,v), now quite non-trivially, solves ∇2 u,vy= 0 in the entire wplane, with y=πon the half-line running from −∞+iπto−1 +iπ, andy=−πon the half-line running from −∞−iπto−1−iπ. We may regard the images of the linesy=const. (solid curves) as being the streamlines of an irrotational and incompressible flow out of the end of a tube into an infinite region, or as the equipotentials near the edge of a pair of capacitor plate s. In the latter case, the images of the lines x=const. (dotted curves) are the corresponding field-lines Example: The Joukowski map . This map is famous in the history of aero- nautics because it can be used to map the exterior of a circle t o the exterior of an aerofoil-shaped region. We can use the Milne-Thomson circle theorem (see 8.3.2) to find the streamlines for the flow past a circle in thezplane, 2Iffwere a function of zonly, then the map would still be isogonal, but would reverse the orientation. We call such maps antiholomorphic oranti-conformal . 8.1. CAUCHY-RIEMANN EQUATIONS 299 -4 -2 2 4 6 -6-4-2246 Figure 8.2: Image of part of the strip −π≤y≤π,−∞<x<∞under the mapz/mapsto→w=z+ez. 300 CHAPTER 8. COMPLEX ANALYSIS I and then use Joukowski’s transformation, w=f(z) =1 2/parenleftbigg z+1 z/parenrightbigg , (8.36) to map this simple flow to the flow past the aerofoil. To produce an aerofoil shape, the circle must go through the point z= 1, where the derivative of f vanishes, and the image of this point becomes the sharp trail ing edge of the aerofoil. The Riemann Mapping Theorem There are tables of conformal maps for D,D/primepairs, but an underlying prin- ciple is provided by the Riemann mapping theorem: Theorem: The interior of any simply connected domain DinCwhose bound- ary consists of more that one point can be mapped conformally one-to-one and onto the interior of the unit circle. It is possible to cho ose an arbitrary interior point w0ofDand map it to the origin, and to take an arbitrary direction through w0and make it the direction of the real axis. With these two choices the mapping is unique . fDw0w Oz Figure 8.3: The Riemann mapping theorem. This theorem was first stated in Riemann’s PhD thesis in 1851. He re- garded it as “obvious” for the reason that we will give as a phy sical “proof.” Riemann’s argument is not rigorous, however, and it was not u ntil 1912 that a real proof was obtained by Constantin Carath´ eodory. A pro of that is both shorter and more in spirit of Riemann’s ideas was given by Leo pold Fej´ er and Frigyes Riesz in 1922. 8.1. CAUCHY-RIEMANN EQUATIONS 301 For the physical “proof,” observe that in the function −1 2πlnz=−1 2π{ln|z|+iθ}, (8.37) the real part φ=−1 2πln|z|is the potential of a unit charge at the origin, and with the additive constant chosen so that φ= 0 on the circle |z|= 1. Now imagine that we have solved the two-dimensional electro statics problem of finding the potential for a unit charge located at w0∈D, also with the boundary of Dbeing held at zero potential. We have ∇2φ1=−δ2(w−w0), φ 1= 0 on∂D. (8.38) Now find the φ2that is harmonically conjugate to φ1. Set φ1+iφ2= Φ(w) =−1 2πln(zeiα) (8.39) whereαis a real constant. We see that the transformation w/mapsto→z, or z=e−iαe−2πΦ(w), (8.40) does the job of mapping the interior of Dinto the interior of the unit circle, and the boundary of Dto the boundary of the unit circle. Note how our freedom to choose the constant αis what allows us to “take an arbitrary direction through w0and make it the direction of the real axis.” Example : To find the map that takes the upper half-plane into the unit circle, with the point z=imapping to the origin, we use the method of images to solve for the complex potential of a unit charge at w=i: φ1+iφ2=−1 2π(ln(w−i)−ln(w+i)) =−1 2πln(eiαz). Therefore z=e−iαw−i w+i. (8.41) We immediately verify that that this works: we have |z|= 1 whenwis real, andz= 0 atw=i. The difficulty with the physical argument is that it is not clea r that a so- lution to the point-charge electrostatics problem exists. In three dimensions, 302 CHAPTER 8. COMPLEX ANALYSIS I for example, there is no solution when the boundary has a shar p inward directed spike. (We cannot physically realize such a situat ion either: the electric field becomes unboundedly large near the tip of a spi ke, and bound- ary charge will leak off and neutralize the point charge.) The re might well be analogous difficulties in two dimensions if the boundary of Dis patho- logical. However, the fact that there isa proof of the Riemann mapping theorem shows that the two-dimensional electrostatics pro blem does always have a solution, at least in the interior ofD— even if the boundary is an infinite-length fractal. However, unless ∂Dis reasonably smooth the result- ing Riemann map cannot be continuously extended to the bound ary. When the boundary of Disa smooth closed curve, then the the boundary of D willmap one-to-one and continuously onto the boundary of the uni t circle. Exercise 8.1 :Van der Pauw’s Theorem.3This problem explains a practical method of for determining the conductivity σof a material, given a sample in the form of of a wafer of uniform thickness d, but of irregular shape. In practice at the Phillips company in Eindhoven, this was a wafer of semi conductor cut from an unmachined boule. A BD C Figure 8.4: A thin semiconductor wafer with attached leads. We attach leads to point contacts A,B,C,D , taken in anticlockwise order, on the periphery of the wafer and drive a current IABfrom A to B. We record the potential difference VD−VCand so find RAB,DC = (VD−VC)/IAB. Similarly we measure RBC,AD . The current flow in the wafer is assumed to be two dimensional, and to obey J=−(σd)∇V,∇·J= 0, 3L. J. Van der Pauw, Phillips Research Reps .13(1958) 1. See also A. M. Thompson, D. G. Lampard, Nature 177(1956) 888, and D. G. Lampard. Proc. Inst. Elec. Eng. C. 104(1957) 271, for the “Calculable Capacitor.” 8.2. COMPLEX INTEGRATION: CAUCHY AND STOKES 303 andn·J= 0 at the boundary (except at the current source and drain). T he potentialVis therefore harmonic, with Neumann boundary conditions. Van der Pauw claims that exp{−πσdRAB,DC}+ exp{−πσdRBC,AD}= 1. From thisσdcan be found numerically. a) First show that Van der Pauw’s claim is true if the wafer wer e the entire upper half-plane with A,B,C,D on the real axis with xA<xB<xC< xD. b) Next, taking care to consider the transformation of the cu rrent source terms and the Neumann boundary conditions, show that the cla im is invariant under conformal maps, and, by mapping the wafer to the upper half-plane, show that it is true in general. 8.2 Complex Integration: Cauchy and Stokes In this section we will define the integral of an analytic func tion, and make contact with the exterior calculus from chapters 2-4. The mo st obvious difference between the real and complex integral is that in ev aluating the definite integral of a function in the complex plane we must sp ecify the path along which we integrate. When this path of integration is th e boundary of a region, it is often called a contour from the use of the word in the graphic arts to describe the outline of something. The integrals the mselves are then called contour integrals . 8.2.1 The Complex Integral The complex integral /integraldisplay Γf(z)dz (8.42) over a path Γ may be defined by expanding out the real and imagin ary parts /integraldisplay Γf(z)dz≡/integraldisplay Γ(u+iv)(dx+idy) =/integraldisplay Γ(udx−vdy)+i/integraldisplay Γ(vdx+udy).(8.43) and treating the two integrals on the right hand side as stand ard vector- calculus line-integrals of the form/integraltext v·dr, one with v→(u,−v) and and one withv→(v,u). 304 CHAPTER 8. COMPLEX ANALYSIS I 0z1 ξ1ξ22zzN N−1ξ N Γzza==ba b Figure 8.5: A chain approximation to the curve Γ. The complex integral can also be constructed as the limit of a Riemann sum in a manner parallel to the definition of the real-variable Ri emann integral of elementary calculus. Replace the path Γ with a chain compo sed of ofN line-segments z0-to-z1,z1-to-z2, all the way to zN−1-to-zN. Now let ξmlie on the line segment joining zm−1andzm. Then the integral/integraltext Γf(z)dzis the limit of the (Riemann) sum N/summationdisplay m=1f(ξm)(zm−zm−1) (8.44) asNgets large and all the |zm−zm−1|→0. For this definition to make sense and be useful, the limit must be independent of both how we chop up the curve and how we select the points ξm. This may be shown to be the case when the integration path is smooth and the function bei ng integrated is continuous. The Riemann-sum definition of the integral leads to a useful i nequality: combining the triangle inequality |a+b|≤|a|+|b|with|ab|=|a||b|we deduce that /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleN/summationdisplay m=1f(ξm)(zm−zm−1)/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤N/summationdisplay m=1|f(ξm)(zm−zm−1)| =N/summationdisplay m=1|f(ξm)||(zm−zm−1)|.(8.45) For sufficiently smooth curves the last sum converges to the re al integral/integraltext Γ|f(z)||dz|, and we deduce that /vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay Γf(z)dz/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/integraldisplay Γ|f(z)||dz|. (8.46) 8.2. COMPLEX INTEGRATION: CAUCHY AND STOKES 305 For curves Γ that are smooth enough to have a well-defined leng th|Γ|, we will have/integraltext Γ|dz|=|Γ|.From this we conclude that if |f|≤Mon Γ, then we have the Darboux inequality /vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay Γf(z)dz/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤M|Γ|. (8.47) We shall find many uses for this inequality. The Riemann sum definition also makes it clear that if f(z) is the deriva- tive of another analytic function g(z),i.e. f(z) =dg dz, (8.48) then, for Γ a smooth path from z=atoz=b, we have /integraldisplay Γf(z)dz=g(b)−g(a). (8.49) This follows by approximating f(ξm)≈(g(zm)−g(zm−1))/(zm−zm−1), and observing that the resultant Riemann sum N/summationdisplay m=1/parenleftBig g(zm)−g(zm−1)/parenrightBig (8.50) telescopes. The approximation to the derivative will becom e exact in the limit|zm−zm−1|→0. Thus, when f(z) is the derivative of another function, the integral is independent of the route that Γ takes from atob. We shall see that any analytic function is (at least locally) the derivative of another analytic function, and so this path independence holds generally — provided that we do not try to move the integration contour o ver a place wherefceases to be differentiable. This is the essence of what is kno wn as Cauchy’s Theorem — although, as with much of complex analysis, the result was known to Gauss. 8.2.2 Cauchy’s theorem Before we state and prove Cauchy’s theorem, we must introduc e an orien- tation convention and some traditional notation. Recall th at ap-chain is a finite formal sum of p-dimensional oriented surfaces or curves, and that a 306 CHAPTER 8. COMPLEX ANALYSIS I p-cycle is ap-chain Γ whose boundary vanishes: ∂Γ = 0. A 1-cycle that con- sists of only a single connected component is a closed curve. We will mostly consider integrals over simple closed curves — these being curves that do not self intersect — or 1-cycles consisting of finite formal sums of such curves. The orientation of a simple closed curve can be described by t he sense, clock- wise or anticlockwise, in which we traverse it. We will adopt the convention that a positively oriented curve is one such that the integra tion is performed in aanticlockwise direction. The integral over a chain Γ of oriented simple closed curves will be denoted by the symbol/contintegraltext Γfdz. We now establish Cauchy’s theorem by relating it to our previ ous work with exterior derivatives: Suppose that fis analytic with a a domain D, so that∂zf= 0 within D. We therefore have that the the exterior derivative of fis df=∂zfdz+∂zfdz=∂zfdz. (8.51) Now suppose that the simple closed curve Γ is the boundary of a region Ω⊂D. We can exploit Stokes’ theorem to deduce that /contintegraldisplay Γ=∂Ωf(z)dz=/integraldisplay Ωd(f(z)dz) =/integraldisplay Ω(∂zf)dz∧dz= 0. (8.52) The last integral is zero because dz∧dz= 0. We may state our result as: Theorem (Cauchy, in modern language): The integral of an ana lytic function over a 1-cycle that is homologous to zero vanishes. The zero result is only guaranteed if the function fis analytic throughout the region Ω. For example, if Γ is the unit circle z=eiθthen /contintegraldisplay Γ/parenleftbigg1 z/parenrightbigg dz=/integraldisplay2π 0e−iθd/parenleftbig eiθ/parenrightbig =i/integraldisplay2π 0dθ= 2πi. (8.53) Cauchy’s theorem is not applicable because 1 /zissingular ,i.e.not differen- tiable, atz= 0. The formula (8.53) will hold for Γ any contour homologous to the unit circle in C\0, the complex plane punctured by the removal of the pointz= 0. Thus/contintegraldisplay Γ/parenleftbigg1 z/parenrightbigg dz= 2πi (8.54) for any contour Γ that encloses the origin. We can deduce a rat her remarkable formula from (8.54): Writing Γ = ∂Ω with anticlockwise orientation, we use Stokes’ theorem to obtain/contintegraldisplay ∂Ω/parenleftbigg1 z/parenrightbigg dz=/integraldisplay Ω∂z/parenleftbigg1 z/parenrightbigg dz∧dz=/braceleftbigg 2πi,0∈Ω, 0,0/∈Ω.(8.55) 8.2. COMPLEX INTEGRATION: CAUCHY AND STOKES 307 Sincedz∧dz= 2idx∧dy, we have established that ∂z/parenleftbigg1 z/parenrightbigg =πδ2(x,y). (8.56) This rather cryptic formula encodes one of the most useful re sults in math- ematics. Perhaps perversely, functions that are more singular than 1 /zhave van- ishing integrals about their singularities. With Γ again th e unit circle, we have/contintegraldisplay Γ/parenleftbigg1 z2/parenrightbigg dz=/integraldisplay2π 0e−2iθd/parenleftbig eiθ/parenrightbig =i/integraldisplay2π 0e−iθdθ= 0. (8.57) The same is true for all higher integer powers: /contintegraldisplay Γ/parenleftbigg1 zn/parenrightbigg dz= 0, n≥2. (8.58) We can understand this vanishing in another way, by evaluati ng the in- tegral as /contintegraldisplay Γ/parenleftbigg1 zn/parenrightbigg dz=/contintegraldisplay Γd dz/parenleftbigg −1 n−11 zn−1/parenrightbigg dz=/bracketleftbigg −1 n−11 zn−1/bracketrightbigg Γ= 0, n/negationslash= 1. (8.59) Here, the notation [ A]Γmeans the difference in the value of Aat two ends of the integration path Γ. For a closed curve the difference is zero because the two ends are at the same point. This approach reinforces t he fact that the complex integral can be computed from the “anti-derivat ive” in the same way as the real-variable integral. We also see why 1 /zis special. It is the derivative of ln z= ln|z|+iargz, and lnzis not really a function, as it is multivalued. In evaluating [ln z]Γwe must follow the continuous evolution of argzas we traverse the contour. As the origin is within the contou r, this angle increases by 2 π, and so [lnz]Γ= [iargz]Γ=i/parenleftbig arge2πi−arge0i/parenrightbig = 2πi. (8.60) Exercise 8.2 : Supposef(z) is analytic in a simply-connected domain D, and z0∈D. Setg(z) =/integraltextz z0f(z)dzalong some path in Dfromz0toz. Use the path-independence of the integral to compute the derivativ e ofg(z) and show that f(z) =dg dz. This confirms our earlier claim that any analytic function is the derivative of some other analytic function. 308 CHAPTER 8. COMPLEX ANALYSIS I Exercise 8.3 :The “D-bar” problem : Suppose we are given a simply-connected domain Ω, and a function f(z,z) defined on it, and wish to find a function F(z,z) such that ∂F(z,z) ∂z=f(z,z),(z,z)∈Ω. Use (8.56) to argue formally that the general solution is F(ζ,¯ζ) =−1 π/integraldisplay Ωf(z,z) z−ζdx∧dy+g(ζ), whereg(ζ) is an arbitrary analytic function. This result can be shown to be correct by more rigorous reasoning. 8.2.3 The residue theorem The essential tool for computations with complex integrals is provided by theresidue theorem . With the aid of this theorem, the evaluation of contour integrals becomes easy. All one has to do is identify points a t which the function being integrated blows up, and examine just how it b lows up. If, near the point zi, the function can be written f(z) =/braceleftBigg a(i) N (z−zi)N+···+a(i) 2 (z−zi)2+a(i) 1 (z−zi)/bracerightBigg g(i)(z), (8.61) whereg(i)(z) is analytic and non-zero at zi, thenf(z) has a poleof orderNat zi. IfN= 1 thenf(z) is said to have a simple pole atzi. We can normalize g(i)(z) so thatg(i)(zi) = 1, and then the coefficient, a(i) 1, of 1/(z−zi) is called the residue of the pole at zi. The coefficients of the more singular terms do not influence the result of the integral, but Nmust be finite for the singularity to be called a pole. Theorem: Let the function f(z)be analytic within and on the boundary Γ =∂Dof a simply connected domain D, with the exception of finite number of points at which f(z)has poles. Then /contintegraldisplay Γf(z)dz=/summationdisplay poles∈D2πi(residue at pole) , (8.62) the integral being traversed in the positive (anticlockwis e) sense . 8.2. COMPLEX INTEGRATION: CAUCHY AND STOKES 309 We prove the residue theorem by drawing small circles Ciabout each singular point ziinD. z3z2 z1 ΓD 1C3C C2 Ω Figure 8.6: Circles for the residue theorem. We now assert that /contintegraldisplay Γf(z)dz=/summationdisplay i/contintegraldisplay Cif(z)dz, (8.63) because the 1-cycle C≡Γ−/summationdisplay iCi=∂Ω (8.64) is the boundary of a region Ω in which fis analytic, and hence Cis homol- ogous to zero. If we make the radius Riof the circle Cisufficiently small, we may replace each g(i)(z) by its limit g(i)(zi) = 1, and so take f(z)→/braceleftBigg a(i) 1 (z−zi)+a(i) 2 (z−zi)2+···+a(i) N (z−zi)N/bracerightBigg g(i)(zi) =a(i) 1 (z−zi)+a(i) 2 (z−zi)2+···+a(i) N (z−zi)N, (8.65) onCi. We then evaluate the integral over Ciby using our previous results to get /contintegraldisplay Cif(z)dz= 2πia(i) 1. (8.66) The integral around Γ is therefore equal to 2 πi/summationtext ia(i) 1. 310 CHAPTER 8. COMPLEX ANALYSIS I The restriction to contours containing only finitely many po les arises for two reasons: Firstly, with infinitely many poles, the sum ove rimight not converge; secondly, there may be a point whose every neighbo urhood contains infinitely many of the poles, and there our construction of dr awing circles around each individual pole would not be possible. Exercise 8.4 :Poisson’s Formula. The function f(z) is analytic in|z|< R/prime. Prove that if|a|<R<R/prime, f(a) =1 2πi/contintegraldisplay |z|=RR2−¯aa (z−a)(R2−¯az)f(z)dz. Deduce that, for 0 <r<R , f(reiθ) =1 2π/integraldisplay2π 0R2−r2 R2−2Rrcos(θ−φ) +r2f(Reiφ)dφ. Show that this formula solves the boundary-value problem fo r Laplace’s equa- tion in the disc|z|<R. Exercise 8.5 :Bergman Kernel. The Hilbert space of analytic functions on a domainDwith inner product /angbracketleftf,g/angbracketright=/integraldisplay D¯fgdxdy is called the Bergman4space ofD. a) Suppose that ϕn(z),n= 0,1,2,..., are a complete set of orthonormal functions on the Bergman space. Show that K(ζ,z) =∞/summationdisplay m=0ϕm(ζ)ϕm(z). has the property that g(ζ) =/integraldisplay/integraldisplay DK(ζ,z)g(z)dxdy. 4This space should not be confused with the Bargmann-Fock spa ce of analytic functions on the entirety of Cwith inner product /angbracketleftf,g/angbracketright=/integraldisplay Ce−|z|2¯fgd2z. Stefan Bergman and Valentine Bargmann are two different peop le. 8.2. COMPLEX INTEGRATION: CAUCHY AND STOKES 311 for any function ganalytic in D. ThusK(ζ,z) plays the role of the delta function on the space of analytic functions on D. This object is called thereproducing orBergman kernel . By taking g(z) =ϕn(z), show that it is the unique integral kernel with the reproducing proper ty. b) Consider the case of Dbeing the unit circle. Use the Gramm-Schmidt procedure to construct an orthonormal set from the function szn,n= 0,1,2,.... Use the result of part a) to conjecture (because we have not proved that the set is complete) that, for the unit circle, K(ζ,z) =1 π1 (1−ζ¯z)2. c) For any smooth, complex valued, function gdefined on a domain Dand its boundary, use Stokes’ theorem to show that /integraldisplay/integraldisplay D∂zg(z,z)dxdy=1 2i/contintegraldisplay ∂Dg(z,z)dz. Use this to verify that this the K(ζ,z) you constructed in part b) is indeed a (and hence “the”) reproducing kernel. d) Now suppose that Dis a simply connected domain whose boundary ∂D is a smooth curve. We know from the Riemann mapping theorem th at there exists an analytic function f(z) =f(z;ζ) that maps Donto the interior of the unit circle in such a way that f(ζ) = 0 andf/prime(ζ) is real and non-zero. Show that if we set K(ζ,z) =f/prime(z)f/prime(ζ)/π, then, by using part c) together with the residue theorem to evaluate the int egral over the boundary, we have g(ζ) =/integraldisplay/integraldisplay DK(ζ,z)g(z)dxdy. ThisK(ζ,z) must therefore be the reproducing kernel. We see that if we knowKwe can recover the map ffrom f/prime(z;ζ) =/radicalbiggπ K(ζ,ζ)K(z,ζ). e) Apply the formula from part d) to the unit circle, and so ded uce that f(z;ζ) =z−ζ 1−¯ζz is the unique function that maps the unit circle onto itself w ith the point ζmapping to the origin and with the horizontal direction thro ughζ remaining horizontal. 312 CHAPTER 8. COMPLEX ANALYSIS I 8.3 Applications We now know enough about complex variables to work through so me inter- esting applications, including the mechanism by which an ae roplane flies. 8.3.1 Two-dimensional vector calculus It is often convenient to use complex co-ordinates for vecto rs and tensors. In these co-ordinates the standard metric on R2becomes “ds2” =dx⊗dx+dy⊗dy =dz⊗dz =gzzdz⊗dz+gzzdz⊗dz+gzzdz⊗dz+gzzdz⊗dz,(8.67) so the complex co-ordinate components of the metric tensor a regzz=gzz= 0, gzz=gzz=1 2. The inverse metric tensor is gzz=gzz= 2,gzz=gzz= 0. In these co-ordinates the Laplacian is ∇2=gij∂2 ij= 2(∂z∂z+∂z∂z). (8.68) Whenfhas singularities, it is not safe to assume that ∂z∂zf=∂z∂zf. For example, from ∂z/parenleftbigg1 z/parenrightbigg =πδ2(x,y), (8.69) we deduce that ∂z∂zlnz=πδ2(x,y). (8.70) When we evaluate the derivatives in the opposite order, howe ver, we have ∂z∂zlnz= 0. (8.71) To understand the source of the non-commutativity, take rea l and imaginary parts of these last two equations. Write ln z= ln|z|+iθ, whereθ= argz, and add and subtract. We find ∇2ln|z|= 2πδ2(x,y), (∂x∂y−∂y∂x)θ= 2πδ2(x,y). (8.72) The first of these shows that1 2πln|z|is the Green function for the Laplace operator, and the second reveals that the vector field ∇θis singular, having a delta function “curl” at the origin. 8.3. APPLICATIONS 313 If we have a vector field vwith contravariant components ( vx,vy) and (nu- merically equal) covariant components ( vx,vy) then the covariant components in the complex co-ordinate system are vz=1 2(vx−ivy) andvz=1 2(vx+ivy). This can be obtained by a using the change of co-ordinates rul e, but a quicker route is to observe that v·dr=vxdx+vydy=vzdz+vzdz. (8.73) Now ∂zvz=1 4(∂xvx+∂yvy) +i1 4(∂yvx−∂xvy). (8.74) Thus the statement that ∂zvz= 0 is equivalent to the vector field vbeing both solenoidal (incompressible) and irrotational. This c an also be expressed in form language by setting η=vzdzand saying that dη= 0 means that the corresponding vector field is both solenoidal and irrotatio nal. 8.3.2 Milne-Thomson Circle Theorem As we mentioned earlier, we can describe an irrotational and incompressible fluid motion either by a velocity potential vx=∂xφ, vy=∂yφ, (8.75) wherevis automatically irrotational but incompressibilty requi res∇2φ= 0, or by a stream function vx=∂yχ, vy=−∂xχ, (8.76) wherevis automatically incompressible but irrotationality requ ires∇2χ= 0. We can combine these into a single complex stream function Φ =φ+iχ which, for an irrotational incompressible flow, satisfies th e Cauchy-Riemann equations and is therefore an analytic function of z. We see that 2vz=dΦ dz, (8.77) φandχmaking equal contributions. The Milne-Thomson theorem says that if Φ is the complex strea m func- tion for a flow in unobstructed space, then /tildewideΦ = Φ(z) +Φ/parenleftbigga2 z/parenrightbigg (8.78) 314 CHAPTER 8. COMPLEX ANALYSIS I is the stream function after the cylindrical obstacle |z|=ais inserted into the flow. Here Φ(z) denotes the analytic function defined by Φ(z) =Φ(z). To see that this works, observe that a2/z=zon the curve|z|=a, and so on this curve Im /tildewideΦ =χ= 0. The surface of the cylinder has therefore become a streamline, and so the flow does not penetrate into the cylin der. If the original flow is created by souces and sinks exterior to |z|=a, which will be singularities of Φ, the additional term has singularites th at lie only within |z|=a. These will be the “images” of the sources and sinks in the sen se of the “method of images.” Example : A uniform flow with speed Uin thexdirection has Φ( z) =Uz. Inserting a cylinder makes this ˜Φ(z) =U/parenleftbigg z+a2 z/parenrightbigg . (8.79) Becausevzis the derivative of this, we see that the perturbing effect of the obstacle on the velocity field falls off as the square of the dis tance from the cylinder. This is a general result for obstructed flows. -2 -1 0 1 2-2-1012 Figure 8.7: The real and imaginary parts of the function z+z−1provide the velocity potentials and streamlines for irrotational inco mpressible flow past a cylinder of unit radius. 8.3.3 Blasius and Kutta-Joukowski Theorems We now derive the celebrated result, discovered independen tly by Martin Wilhelm Kutta (1902) and Nikolai Egorovich Joukowski (1906 ), that the 8.3. APPLICATIONS 315 lift per unit span of an aircraft wing is equal to the product o f the density of the airρ, the circulation κ≡/contintegraltext v·drabout the wing, and the forward velocityUof the wing through the air. Their theory treats the air as bei ng incompressible—a good approximation unless the flow-veloc ities approach the speed of sound—and assumes that the wing is long enough th at the flow can be regarded as being two dimensional. UF Figure 8.8: Flow past an aerofoil. Begin by recalling how the momentum flux tensor Tij=ρvivj+gijP (8.80) enters fluid mechanics. In cartesian co-ordinates, and in th e presence of an external body force fiacting on the fluid, the Euler equation of motion for the fluid is ρ(∂tvi+vj∂jvi) =−∂iP+fi. (8.81) HerePis the pressure and we are distinguishing between co and cont ravariant components, although at the moment gij≡δij. We can combine Euler’s equation with the law of mass conservation, ∂tρ+∂i(ρvi) = 0, (8.82) to obtain ∂t(ρvi) +∂j(ρvjvi+gijP) =fi. (8.83) This momemtum-tracking equation shows that the external fo rce acts as a source of momentum, and that for steady flow fiis equal to the divergence of the momentum flux tensor: fi=∂lTli=gkl∂kTli. (8.84) 316 CHAPTER 8. COMPLEX ANALYSIS I As we are interested in steady, irrotational motion with uni form density we may use Bernoulli’s theorem, P+1 2ρ|v|2=const. , to substitute−1 2ρ|v|2in place ofP. (The constant will not affect the momentum flux.) With this substitution Tijbecomes a traceless symmetric tensor: Tij=ρ(vivj−1 2gij|v|2). (8.85) Usingvz=1 2(vx−ivy) and Tzz=∂xi ∂z∂xj ∂zTij, (8.86) together with x≡x1=1 2(z+z), y≡x2=1 2i(z−z) (8.87) we find T≡Tzz=1 4(Txx−Tyy−2iTxy) =ρ(vz)2. (8.88) This is the only component of Tijthat we will need to consider. Tzzis simply T, whereasTzz= 0 =TzzbecauseTijis traceless. In our complex co-ordinates, the equation fi=gkl∂kTli (8.89) reads fz=gzz∂zTzz+gzz∂zTzz= 2∂zT. (8.90) We see that in steady flow the net momentum flux ˙Piout of a region Ω is given by ˙Pz=/integraldisplay Ωfzdxdy=1 2i/integraldisplay Ωfzdzdz=1 i/integraldisplay Ω∂zTdzdz=1 i/contintegraldisplay ∂ΩTdz. (8.91) We have used Stokes’ theorem at the last step. In regions wher e there is no external force, Tis analytic, ∂zT= 0, and the integral will be independent of the choice of contour ∂Ω. We can subsititute T=ρv2 zto get ˙Pz=−iρ/contintegraldisplay ∂Ωv2 zdz. (8.92) 8.3. APPLICATIONS 317 To apply this result to our aerofoil we take can take ∂Ω to be its boundary. Then ˙Pzis the total force exerted on the fluid by the wing, and, by Newt on’s third law, this is minus the force exerted by the fluid on the wi ng. The total force on the aerofoil is therefore Fz=iρ/contintegraldisplay ∂Ωv2 zdz. (8.93) The result (8.93) is often called Blasius’ theorem . Evaluating the integral in (8.93) is not immediately possib le because the velocity von the boundary will be a complicated function of the shape of the body. We can, however, exploit the contour independence of the integral and evaluate it over a path encircling the aerofoil at large d istance where the flow field takes the asymptotic form vz=Uz+κ 4πi1 z+O/parenleftbigg1 z2/parenrightbigg . (8.94) TheO(1/z2) term is the velocity perturbation due to the air having to flo w round the wing, as with the cylinder in a free flow. To confirm th at this flow has the correct circulation we compute /contintegraldisplay v·dr=/contintegraldisplay vzdz+/contintegraldisplay vzdz=κ. (8.95) Substituting vzin (8.93) we find that the O(1/z2) term cannot contribute as it cannot affect the residue of any pole. The only part that doe s contribute is the cross term that arises from multiplying Uzbyκ/(4πiz). This gives Fz=iρ/parenleftbiggUzκ 2πi/parenrightbigg/contintegraldisplaydz z=iρκUz (8.96) so that1 2(Fx−iFy) =iρκ1 2(Ux−iUy). (8.97) Thus, in conventional co-ordinates, the reaction force on t he body is Fx=ρκUy, Fy=−ρκUx. (8.98) The fluid therefore provides a lift force proportional to the product of the circulation with the asymptotic velocity. The force is at ri ght angles to the incident airstream, so there is no drag. 318 CHAPTER 8. COMPLEX ANALYSIS I The circulation around the wing is determined by the Kutta condition that the velocity of the flow at the sharp trailing edge of the w ing be finite. If the wing starts moving into the air and the requisite circu lation is not yet established then the flow under the wing does not leave the trailing edge smoothly but tries to whip round to the topside. The velocity gradients become very large and viscous forces become important and pr event the air from making the sharp turn. Instead, a starting vortex is shed from the trailing edge. Kelvin’s theorem on the conservation of vort icity shows that this causes a circulation of equal and opposite strength to b e induced about the wing. For finite wings, the path independence of/contintegraltext v·drmeans that the wings must leave a pair of trailing wingtip vortices of strength κthat connect back to the starting vortex to form a closed loop. The velocity fiel d induced by the trailing vortices cause the airstream incident on the aerof oil to come from a slighly different direction than the asymptotic flow. Conseq uently, the lift is not quite perpendicular to the motion of the wing. For finite- length wings, therefore, lift comes at the expense of an inevitable induced drag force. The work that has to be done against this drag force in driving the wing forwards provides the kinetic energy in the trailing vortices. 8.4 Applications of Cauchy’s Theorem Cauchy’s theorem provides the Royal Road to complex analysi s. It is possible to develop the theory without it, but the path is harder going . 8.4.1 Cauchy’s Integral Formula Iff(z) is analytic within and on the boundary of a simply connected domain Ω, with∂Ω = Γ, and if ζis a point in Ω, then, noting that the the integrand has a simple pole at z=ζand applying the residue formula, we have Cauchy’s integral formula f(ζ) =1 2πi/contintegraldisplay Γf(z) z−ζdz, ζ∈Ω. (8.99) 8.4. APPLICATIONS OF CAUCHY’S THEOREM 319 Γ ζΩ Figure 8.9: Cauchy contour. This formula holds only if ζlies within Ω. If it lies outside, then the integrand is analytic everywhere inside Ω, and so the integral gives ze ro. We may show that it is legitimate to differentiate under the in tegral sign in Cauchy’s formula. If we do so ntimes, we have the useful corollary that f(n)(ζ) =n! 2πi/contintegraldisplay Γf(z) (z−ζ)n+1dz. (8.100) This shows that being oncedifferentiable (analytic) in a region automatically implies that f(z) is differentiable arbitrarily many times ! Exercise 8.6 :The generalized Cauchy formula . Suppose that we have solved a D-bar problem (see exercise 8.3), and so found an F(z,z) with∂zF=f(z,z) in a region Ω. Compute the exterior derivative of F(z,z) z−ζ using (8.56). Now, manipulating formally with delta functi ons, apply Stokes’ theorem to show that, for ( ζ,¯ζ) in the interior of Ω, we have F(ζ,¯ζ) =1 2πi/contintegraldisplay ∂ΩF(z,z) z−ζdz−1 π/integraldisplay Ωf(z,z) z−ζdxdy. This is called the generalized Cauchy formula . Note that the first term on the right, unlike the second, is a function only of ζ, and so is analytic. Liouville’s Theorem A dramatic corollary of Cauchy’s integral formula is provid ed by 320 CHAPTER 8. COMPLEX ANALYSIS I Liouville’s theorem :Iff(z)is analytic in all of C, and is bounded there, meaning that there is a positive real number Ksuch that|f(z)|<K, then f(z)is a constant. This result provides a powerful strategy for proving that tw o formulæ, f1(z) andf2(z), represent the same analytic function. If we can show that the difference f1−f2is analytic and tends to zero at infinity then Liouville’s theorem tells us that f1=f2. Because the result is perhaps unintuitive, and because the m ethods are typical, we will spell out in detail how Liouville’s theorem works. We select any two points, z1andz2, and use Cauchy’s formula to write f(z1)−f(z2) =1 2πi/contintegraldisplay Γ/parenleftbigg1 z−z1−1 z−z2/parenrightbigg f(z)dz. (8.101) We take the contour Γ to be circle of radius ρcentered on z1. We make ρ>2|z1−z2|, so that when zis on Γ we are sure that |z−z2|>ρ/2. >ρ/2 ρz2 z1z Figure 8.10: Contour for Liouville’ theorem. Then, using|/integraltext f(z)dz|≤/integraltext |f(z)||dz|, we have |f(z1)−f(z2)|=1 2π/vextendsingle/vextendsingle/vextendsingle/vextendsingle/contintegraldisplay Γ(z1−z2) (z−z1)(z−z2)f(z)dz/vextendsingle/vextendsingle/vextendsingle/vextendsingle ≤1 2π/integraldisplay2π 0|z1−z2|K ρ/2dθ=2|z1−z2|K ρ.(8.102) The right hand side can be made arbitrarily small by taking ρlarge enough, so we we must have f(z1) =f(z2). Asz1andz2were any pair of points, we deduce that f(z) takes the same value everywhere. 8.4. APPLICATIONS OF CAUCHY’S THEOREM 321 8.4.2 Taylor and Laurent Series We have defined a function to be analytic in a domain Dif it is (once) complex differentiable at all points in D. It turned out that this apparently mild requirement automatically implied that the function i s differentiable arbitrarily many times inD. In this section we shall see that knowledge of all derivatives of f(z) at any single point in Dis enough to completely determine the function at any other point in D. Compare this with functions of a real variable, for which it is easy to construct examples that are once but not twice differentiable, and where complete knowledge o f function at a point, or in even in a neighbourhood of a point, tells us absol utely nothing of the behaviour of the function away from the point or neighb ourhood. The key ingredient in these almost magical properties of com plex ana- lytic functions is that any analytic function has a Taylor se ries expansion that actually converges to the function. Indeed an alternat ive definition of analyticity is that f(z) be representable by a convergent power series. For real variables this is the definition of a real analytic function. To appreciate the utility of power series representations w e do need to discuss some basic properties of power series. Most of these results are ex- tensions to the complex plane of what we hope are familiar not ions from real analysis. Consider the power series ∞/summationdisplay n=0an(z−z0)n≡lim N→∞SN, (8.103) whereSNare the partial sums SN=N/summationdisplay n=0an(z−z0)n. (8.104) Suppose that this limit exists (i.e the series is convergent ) for some z=ζ; then it turns out that the series is absolutely convergent5for any|z−z0|< |ζ−z0|. 5Recall that absolute convergence of/summationtextanmeans that/summationtext|an|converges. Absolute convergence implies convergence, and also allows us to rear range the order of terms in the series without changing the value of the sum. Compare this wi thconditional convergence , where/summationtextanconverges, but/summationtext|an|does not. You may remember that Riemann showed that the terms of a conditionally convergent series can be re arranged so as to get any answer whatsoever ! 322 CHAPTER 8. COMPLEX ANALYSIS I To establish this absolute convergence we may assume, witho ut loss of generality, that z0= 0. Then, convergence of the sum/summationtextanζnrequires that |anζn|→0, and thus|anζn|is bounded. In other words, there is a Bsuch that|anζn|<Bfor anyn. We now write |anzn|=|anζn|/vextendsingle/vextendsingle/vextendsingle/vextendsinglez ζ/vextendsingle/vextendsingle/vextendsingle/vextendsinglen <B/vextendsingle/vextendsingle/vextendsingle/vextendsinglez ζ/vextendsingle/vextendsingle/vextendsingle/vextendsinglen . (8.105) The sum/summationtext|anzn|therefore converges for |z/ζ|<1, by comparison with a geometric progression. This result, that if a power series in ( z−z0) converges at a point then it converges at all points closer to z0, shows that a power series possesses someradius of convergence R. The series converges for all |z−z0|<R, and diverges for all|z−z0|> R. (What happens onthe circle|z−z0|=Ris usually delicate, and harder to establish.) We soon show tha t the radius of convergence of a power series is the distance from z0to the nearest singularity of the function that it represents. By comparison with a geometric progression, we may establis h the fol- lowing useful formulæ giving Rfor the series/summationtextanzn: R= lim n→∞|an−1| |an| = lim n→∞|an|1/n. (8.106) The proof of these formulæ is identical the real-variable ve rsion. When we differentiate the terms in a power series, and thus tak eanzn→ nanzn−1, this does not alter R. This observation suggests that it is legitimate to evaluate the derivative of the function represented by th e powers series by differentiating term-by-term. As step on the way to justifyi ng this, observe that if the series converges at z=ζandDris the domain|z|<r<|ζ|then, using the same bound as in the proof of absolute convergence, we have |anzn|<B|zn| |ζ|n<Brn |ζ|n=Mn (8.107) where/summationtextMnis convergent. As a consequence/summationtextanznisuniformly con- vergent inDrby the Weierstrass “ M” test. You probably know that uni- form convergence allows the interchange the order of sums an dintegrals :/integraltext (/summationtextfn(x))dx=/summationtext/integraltext fn(x)dx. For real variables, uniform convergence is 8.4. APPLICATIONS OF CAUCHY’S THEOREM 323 nota strong enough a condition for us to to safely interchange or der of sums andderivatives : (/summationtextfn(x))/primeis not necessarily equal to/summationtextf/prime n(x). For complex analytic functions, however, Cauchy’s integral formula re duces the operation of differentiation to that of integration, and so this interc hange ispermitted. In particular we have that if f(z) =∞/summationdisplay n=0anzn, (8.108) andRis defined by R=|ζ|for anyζfor which the series converges, then f(z) is analytic in|z|<Rand f/prime(z) =∞/summationdisplay n=0nanzn−1, (8.109) is also analytic in |z|<R. Morera’s Theorem There is is a partial converse of Cauchy’s theorem: Theorem (Morera): If f(z)is defined and continuous in a domain D, and if/contintegraltext Γf(z)dz= 0for all closed contours, then f(z)is analytic in D.To prove this we set F(z) =/integraltextz Pf(ζ)dζ. The integral is path-independent by the hypothesis of the theorem, and because f(z) is continuous we can differentiate with respect to the integration limit to find that F/prime(z) =f(z). ThusF(z) is complex differentiable, and so analytic. Then, by Cauchy’ s formula for higher derivatives, F/prime/prime(z) =f/prime(z) exists, and so f(z) itself is analytic. A corollary of Morera’s theorem is that if fn(z)→f(z) uniformly in D, with all the fnanalytic, then i)f(z) is analytic in D, and ii)f/prime n(z)→f/prime(z) uniformly. We use Morera’s theorem to prove (i) (appealing to the unifor m conver- gence to justify the interchange the order of summation and i ntegration), and use Cauchy’s theorem to prove (ii). Taylor’s Theorem for analytic functions Theorem: Let Γbe a circle of radius ρcentered on the point a. Suppose that f(z)is analytic within and on Γ, and and that the point z=ζis within Γ. 324 CHAPTER 8. COMPLEX ANALYSIS I Thenf(ζ)can be expanded as a Taylor series f(ζ) =f(a) +∞/summationdisplay n=1(ζ−a)n n!f(n)(a), (8.110) meaning that this series converges to f(ζ)for allζsuch that|ζ−a|<ρ. To prove this theorem we use identity 1 z−ζ=1 z−a+(ζ−a) (z−a)2+···+(ζ−a)N−1 (z−a)N+(ζ−a)N (z−a)N1 z−ζ(8.111) and Cauchy’s integral, to write f(ζ) =1 2πi/contintegraldisplay Γf(z) (z−ζ)dz =N−1/summationdisplay n=0(ζ−a)n 2πi/contintegraldisplayf(z) (z−a)n+1dz+(ζ−a)N 2πi/contintegraldisplayf(z) (z−a)N(z−ζ)dz =N−1/summationdisplay n=0(ζ−a)n n!f(n)(a) +RN, (8.112) where RNdef=(ζ−a)N 2πi/contintegraldisplay Γf(z) (z−a)N(z−ζ)dz. (8.113) This is Taylor’s theorem with remainder. For real variables this is as far as we can go. Even if a real function is differentiable infinitely many times, there is no reason for the remainder to become small. For anal ytic functions, however, we can show that RN→0 asN→ ∞ . This means that the complex-variable Taylor series is convergent, and its limi t is actually equal tof(z). To show that RN→0, recall that Γ is a circle of radius ρcentered onz=a. Letr=|ζ−a|<ρ, and letMbe an upper bound for f(z) on Γ. (This exists because fis continuous and Γ is a compact subset of C.) Then, estimating the integral using methods similar to those invo ked in our proof of Liouville’s Theorem, we find that RN<rN 2π/parenleftbigg2πρM ρN(ρ−r)/parenrightbigg . (8.114) Asr<ρ, this tends to zero as N→∞. 8.4. APPLICATIONS OF CAUCHY’S THEOREM 325 We can take ρas large as we like provided there are no singularities of fend up within, or on, the circle. This confirms the claim made e arlier: the radius of convergence of the powers series representati on of an analytic functionis the distance to the nearest singularity. Laurent Series Theorem (Laurent): Let Γ1andΓ2be two anticlockwise circlular paths with centrea, radiiρ1andρ2, and withρ2<ρ1. Iff(z)is analytic on the circles and within the annulus between them, then, for ζin the annulus : f(ζ) =∞/summationdisplay n=0an(ζ−a)n+∞/summationdisplay n=1bn(ζ−a)−n. (8.115) Γ1Γ2 ζ a Figure 8.11: Contours for Laurent’s theorem. The coefficients anandbnare given by an=1 2πi/contintegraldisplay Γ1f(z) (z−a)n+1dz, bn=1 2πi/contintegraldisplay Γ2f(z)(z−a)n−1dz. (8.116) Laurent’s theorem is proved by observing that f(ζ) =1 2πi/contintegraldisplay Γ1f(z) (z−ζ)dz−1 2πi/contintegraldisplay Γ2f(z) (z−ζ)dz, (8.117) and using the identities 1 z−ζ=1 z−a+(ζ−a) (z−a)2+···+(ζ−a)N−1 (z−a)N+(ζ−a)N (z−a)N1 z−ζ,(8.118) 326 CHAPTER 8. COMPLEX ANALYSIS I and −1 z−ζ=1 ζ−a+(z−a) (ζ−a)2+···+(z−a)N−1 (ζ−a)N+(z−a)N (ζ−a)N1 ζ−z.(8.119) Once again we can show that the remainder terms tend to zero. Warning : Although the coefficients anare given by the same integrals as in Taylor’s theorem, they are not interpretable as derivative s offunlessf(z) is analytic within the inner circle, in which case all the bnare zero. 8.4.3 Zeros and Singularities This section is something of a nosology — a classification of diseases — but you should study it carefully as there is some tight reasonin g here, and the conclusions are the essential foundations for the rest of su bject. First a review and some definitions: a) Iff(z) is analytic with a domain D, we have seen that fmay be expanded in a Taylor series about any point z0∈D: f(z) =∞/summationdisplay n=0an(z−z0)n. (8.120) Ifa0=a1=···=an−1= 0, andan/negationslash= 0, so that the first non-zero term in the series is an(z−z0)n, we say that f(z) has a zeroof ordern atz0. b) Asingularity off(z) is a point at which f(z) ceases to be differentiable. Iff(z) has no singularities at finite z(for example, f(z) = sinz) then it is said to be an entire function. c) Iff(z) is analytic in Dexcept atz=a, anisolated singularity , then we may draw two concentric circles of centre a, both within D, and in the annulus between them we have the Laurent expansion f(z) =∞/summationdisplay n=0an(z−a)n+∞/summationdisplay n=1bn(z−a)−n. (8.121) The second term, consisting of negative powers, is called th eprincipal partoff(z) atz=a. It may happen that bm/negationslash= 0 butbn= 0,n>m . Such a singularity is called a pole of order matz=a. The coefficient b1, which may be 0, is called the residue of fat the pole z=a. If the series of negative powers does not terminate, the singulari ty is called anisolated essential singularity 8.4. APPLICATIONS OF CAUCHY’S THEOREM 327 Now some observations: i) Suppose f(z) is analytic in a domain Dcontaining the point z=a. Then we can expand: f(z) =/summationtextan(z−a)n. Iff(z) is zero at z= 0, then there are exactly two possibilities: a) all the anvanish, and then f(z) is identically zero; b) there is a first non-zero coefficient, amsay, and sof(z) =zmϕ(z), whereϕ(a)/negationslash= 0. In the second case fis said to possess a zero of order matz=a. ii) Ifz=ais a zero of order m, off(z) then the zero is isolated –i.e. there is a neighbourhood of awhich contains no other zero. To see this observe that f(z) = (z−a)mϕ(z) whereϕ(z) is analytic and ϕ(a)/negationslash= 0. Analyticity implies continuity, and by continuity there is a neighbour- hood ofain whichϕ(z) does not vanish. iii) Limit points of zeros I: Suppose that we know that f(z) is analytic in D and we know that it vanishes at a sequence of points a1,a2,a3,...∈D. If these points have a limit point6that is interior to Dthenf(z) must, by continuity, be zero there. But this would be a non-isolate d zero, in contradiction to item ii), unless f(z) actually vanishes identically in D. This, then, is the only option. iv) From the definition of poles, they too are isolated. v) Iff(z) has a pole at z=athenf(z)→∞ asz→ain any manner. vi) Limit points of zeros II: Suppose we know that fis analytic in D, except possibly at z=awhich is limit point of zeros as in iii), but we also know that fis not identically zero. Then z=amust be singularity off— but not a pole ( because fwould tend to infinity and could not have arbitrarily close zeros) — so amust be an isolated essential singularity. For example sin 1 /zhas an isolated essential singularity at z= 0, this being a limit point of the zeros at z= 1/nπ. vii) A limit point of poles or other singularities would be a non-isolated essential singularity . 8.4.4 Analytic Continuation Suppose that f1(z) is analytic in the (open, arcwise-connected) domain D1, andf2(z) is analytic in D2, withD1∩D2/negationslash=∅. Suppose further that f1(z) = f2(z) inD1∩D2. Then we say that f2is an analytic continuation of f1to 6A pointz0is a limit point of a set Sif for every /epsilon1>0 there is some a∈S, other than z0itself, such that|a−z0|≤/epsilon1. A sequence need not have a limit for it to possess one or more limit points. 328 CHAPTER 8. COMPLEX ANALYSIS I D2. Such analytic continuations are unique : iff3is also analytic in D2, and f3=f1inD1∩D2, thenf2−f3= 0 inD1∩D2. Because the intersection of two open sets is also open, f1−f2vanishes on an open set and, so by observation iii) of the previous section, it vanishes every where inD2. D1D2 Figure 8.12: Intersecting domains. We can use this uniqueness result, coupled with the circular domains of convergence of the Taylor series, to extend the definition of analytic functions beyond the domain of their initial definition. The distribution xα−1 + An interesting and useful example of analytic continuation is provided by the distribution xα−1 +, which, for real positive α, is defined by its evaluation on a test function ϕ(x) as (xα−1 +,ϕ) =/integraldisplay∞ 0xα−1ϕ(x)dx. (8.122) The pairing ( xα−1 +,ϕ) extends to an complex analytic function of αprovided the integral converges. Test functions are required to decr ease at infinity faster than any power of x, and so the integral always converges at the upper limit. It will converge at the lower limit provided Re ( α)>0. Assume that this is so, and integrate by parts using d dx/parenleftbiggxα αϕ(x)/parenrightbigg =xα−1ϕ(x) +xα αϕ/prime(x). (8.123) We find that, for /epsilon1>0, /bracketleftbiggxα αϕ(x)/bracketrightbigg∞ /epsilon1=/integraldisplay∞ /epsilon1xα−1ϕ(x)dx+/integraldisplay∞ /epsilon1xα αϕ/prime(x)dx. (8.124) 8.4. APPLICATIONS OF CAUCHY’S THEOREM 329 The integrated-out part on the left-hand-side of (8.124) te nds to zero as we take/epsilon1to zero, and both of the integrals converge in this limit as we ll. Consequently I1(α)≡−1 α/integraldisplay∞ 0xαϕ/prime(x)dx (8.125) is equal to ( xα−1 +,ϕ) for 0<Re (α)<∞. However, the integral defining I1(α) converges in the larger region −1<Re (α)<∞. It therefore provides an analytic continuation to this larger domain. The factor o f 1/αreveals that the analytically-continued function possesses a pole at α= 0, with residue −/integraldisplay∞ 0ϕ/prime(x)dx=ϕ(0). (8.126) We can repeat the integration by parts, and find that I2(α)≡1 α(α+ 1)/integraldisplay∞ 0xα+1ϕ/prime/prime(x)dx (8.127) provides an analytic continuation to the region −2<Re(α)<∞. By proceeding in this manner, we can continue ( xα−1 +,ϕ) to a function analytic in the entire complex αplane with the exception of zero and the negative integers, at which it has simple poles. The residue of the pol e atα=−nis ϕ(n)(0)/n!. There is another, much more revealing, way of expressing the se analytic continuations. To obtain this, suppose that φ∈C∞[0,∞] andφ→0 at infinity as least as fast as 1 /x. (Our test function ϕdecreases much more rapidly than this, but 1 /xis all we need for what follows.) Now the function I(α)≡/integraldisplay∞ 0xα−1φ(x)dx (8.128) is convergent and analytic in the strip 0 <Re (α)<1. By the same reasoning as above,I(α) is there equal to −/integraldisplay∞ 0xα αφ/prime(x)dx. (8.129) Again this new integral provides an analytic continuation t o the larger strip −1<Re (α)<1. But in the left-hand half of this strip, where −1< 330 CHAPTER 8. COMPLEX ANALYSIS I Re(α)<0, we can write −/integraldisplay∞ 0xα αφ/prime(x)dx= lim /epsilon1→0/braceleftbigg/integraldisplay∞ /epsilon1xα−1φ(x)dx−/bracketleftbiggxα αφ(x)/bracketrightbigg∞ /epsilon1/bracerightbigg = lim /epsilon1→0/braceleftbigg/integraldisplay∞ /epsilon1xα−1φ(x)dx+φ(/epsilon1)/epsilon1α α/bracerightbigg = lim /epsilon1→0/braceleftbigg/integraldisplay∞ /epsilon1xα−1[φ(x)−φ(/epsilon1)]dx/bracerightbigg , =/integraldisplay∞ 0xα−1[φ(x)−φ(0)]dx. (8.130) Observe how the integrated out part, which tends to zero in 0 <Re (α)<1, becomes divergent in the strip −1<Re (α)<0. This divergence is there craftily combined with the integral to cancel itsdivergence, leaving a finite remainder. As a consequence, for −1<Re (α)<0, the analytic continuation is given by I(α) =/integraldisplay∞ 0xα−1[φ(x)−φ(0)]dx. (8.131) Next we observe that χ(x) = [φ(x)−φ(0)]/xtends to zero as 1 /xfor largex, and atx= 0 can be defined by its limit as χ(0) =φ/prime(0). Thisχ(x) then satisfies the same hypotheses as φ(x). WithI(α) denoting the analytic continuation of the original I, we therefore have I(α) =/integraldisplay∞ 0xα−1[φ(x)−φ(0)]dx,−1<Re (α)<0 =/integraldisplay∞ 0xβ−1/bracketleftbiggφ(x)−φ(0) x/bracketrightbigg dx, whereβ=α+ 1, →/integraldisplay∞ 0xβ−1/bracketleftbiggφ(x)−φ(0) x−φ/prime(0)/bracketrightbigg dx,−1<Re(β)<0 =/integraldisplay∞ 0xα−1[φ(x)−φ(0)−xφ/prime(0)]dx,−2<Re (α)<−1, (8.132) the arrow denoting the same analytic continuation process t hat we used with φ. We can now apply this machinary to our original ϕ(x), and so deduce 8.4. APPLICATIONS OF CAUCHY’S THEOREM 331 that the analytically-continued distribution is given by (xα−1 +,ϕ) =  /integraldisplay∞ 0xα−1ϕ(x)dx, 0<Re (α)<∞, /integraldisplay∞ 0xα−1[ϕ(x)−ϕ(0)]dx,−1<Re (α)<0, /integraldisplay∞ 0xα−1[ϕ(x)−ϕ(0)−xϕ/prime(0)]dx,−2<Re (α)<−1, (8.133) and so on. The analytic continuation automatically subtrac ts more and more terms of the Taylor series of ϕ(x) the deeper we penetrate into the left-hand half-plane. This property, that analytic continuation cov ertly subtracts the minimal number of Taylor-series terms required ensure conv ergence, lies be- hind a number of physics applications, most notably the meth od ofdimen- sional regularization in quantum field theory. The following exercise illustrates some standard techniqu es of reasoning viaanalytic continuation. Exercise 8.7 : Define the dilogarithm function by the series Li2(z) =z 12+z2 22+z3 32+···. The radius of convergence of this series is unity, but the dom ain of Li 2(z) can be extended to|z|>1 by analytic continuation. a) Observe that the series converges at z=±1, and atz= 1 is Li2(1) = 1 +1 22+1 32+···=π2 6. Rearrange the series to show that Li2(−1) =−π2 12. b) Identify the derivative of the power series for Li 2(z) with that of an elementary function. Exploit your identification to extend the definition of [Li 2(z)]/primeoutside|z|<1. Use the properties of this derivative function, together with part a), to prove that Li2(−z) + Li 2/parenleftbigg −1 z/parenrightbigg =−1 2(lnz)2−π2 6. This formula allows us to calculate values of the dilogarith m for|z|>1 in terms of those with |z|<1. 332 CHAPTER 8. COMPLEX ANALYSIS I Many weird identities involving dilogarithms exist. Some, such as Li2/parenleftbigg −1 2/parenrightbigg +1 6Li2/parenleftbigg1 9/parenrightbigg =−1 18π2+ ln 2ln 3−1 2(ln 2)2−1 3(ln 3)2, were found by Ramanujan. Others, originally discovered by s ophisticated numerical methods, have been given proofs based on techniqu es from quantum mechanics. Polylogarithms , defined by Lik(z) =z 1k+z2 2k+z3 3k+···, occur frequently when evaluating Feynman diagrams. 8.4.5 Removable Singularities and the Weierstrass-Casora ti Theorem Sometimes we are given a definition that makes a function anal ytic in a region with the exception of a single point. Can we extend the definition to make the function analytic in the entire region? Provided th at the function is well enough behaved near the point, the answer is yes, and t he extension is unique. Curiously, the proof that this is so gives us insig ht into the wild behaviour of functions near essential singularities. Removable singularities Suppose that f(z) is analytic in D\a, but that lim z→a(z−a)f(z) = 0, then f may be extended to a function analytic in all of D—i.e.z=ais aremovable singularity . To see this, let ζlie between two simple closed contours Γ 1and Γ2, withawithin the smaller, Γ 2. We use Cauchy to write f(ζ) =1 2πi/contintegraldisplay Γ1f(z) z−ζdz−1 2πi/contintegraldisplay Γ2f(z) z−ζdz. (8.134) Now we can shrink Γ 2down to be very close to a, and because of the condition onf(z) nearz=a, we see that the second integral vanishes. We can also arrange for Γ 1to enclose any chosen point in D. Thus, if we set ˜f(ζ) =1 2πi/contintegraldisplay Γ1f(z) z−ζdz (8.135) within Γ 1, we see that ˜f=finD\a, and is analytic in all of D. The extension is unique because any two analytic functions that agree ever ywhere except for a single point, must also agree at that point. 8.4. APPLICATIONS OF CAUCHY’S THEOREM 333 Weierstrass-Casorati We apply the idea of removable singularities to show just how pathological a beast is an isolated essential singularity: Theorem (Weierstrass-Casorati): Let z=abe an isolated essential singular- ity off(z), then in any neighbourhood of athe function f(z)comes arbitrarily close to any assigned valued in C. To prove this, define Nδ(a) ={z∈C:|z−a|< δ}, andN/epsilon1(ζ) ={z∈ C:|z−ζ|< /epsilon1}. The claim is then that there is an z∈Nδ(a) such that f(z)∈N/epsilon1(ζ). Suppose that the claim is nottrue. Then we have |f(z)−ζ|>/epsilon1 for allz∈Nδ(a). Therefore /vextendsingle/vextendsingle/vextendsingle/vextendsingle1 f(z)−ζ/vextendsingle/vextendsingle/vextendsingle/vextendsingle<1 /epsilon1(8.136) inNδ(a), while 1/(f(z)−ζ) is analytic in Nδ(a)\a. Therefore z=ais a removable singularity of 1 /(f(z)−ζ), and there is an an analytic g(z) which coincides with 1 /(f(z)−ζ) at all points except a. Therefore f(z) =ζ+1 g(z)(8.137) except ata. Nowg(z), being analytic, may have a zero at z=agiving a pole inf, but it cannot give rise to an essential singularity. The cla im is true, therefore. Picard’s Theorems Weierstrass-Casorati is elementary. There are much strong er results: Theorem (Picard’s little theorem): Every nonconstant enti re function attains every complex value with at most oneexception. Theorem (Picard’s big theorem): In any neighbourhood of an i solated essen- tial singularity, f(z)takes every complex value with at most oneexception. The proofs of these theorems are hard. As an illustration of Picard’s little theorem, observe that the function expzis entire, and takes all values except 0. For the big theorem o bserve that function f(z) = exp(1/z). has an essential singularity at z= 0, and takes all values, with the exception of 0, in any neighbourho od ofz= 0. 334 CHAPTER 8. COMPLEX ANALYSIS I 8.5 Meromorphic functions and the Winding- Number A function whose only singularities in Dare poles is said to be meromor- phicthere. These functions have a number of properties that are e ssentially topological in character. 8.5.1 Principle of the Argument Iff(z) is meromorphic in Dwith∂D= Γ, andf(z)/negationslash= 0 on Γ, then 1 2πi/contintegraldisplay Γf/prime(z) f(z)dz=N−P (8.138) whereNis the number of zero’s in DandPis the number of poles. To show this, we note that if f(z) = (z−a)mϕ(z) whereϕis analytic and non-zero neara, then f/prime(z) f(z)=m z−a+ϕ/prime(z) ϕ(z)(8.139) sof/prime/fhas a simple pole at awith residue m. Heremcan be either positive or negative. The term ϕ/prime(z)/ϕ(z) is analytic at z=a, so collecting all the residues from each zero or pole gives the result. Sincef/prime/f=d dzlnfthe integral may be written /contintegraldisplay Γf/prime(z) f(z)dz= ∆ Γlnf(z) =i∆Γargf(z), (8.140) the symbol ∆ Γdenoting the total change in the quantity after we traverse Γ . Thus N−P=1 2π∆Γargf(z). (8.141) This result is known as the principle of the argument. Local mapping theorem Suppose the function w=f(z) maps a region Ω holomorphicly onto a region Ω/prime, and a simple closed curve γ⊂Ω onto another closed curve Γ ⊂Ω/prime, which will in general have self intersections. Given a point a∈Ω/prime, we can ask 8.5. MEROMORPHIC FUNCTIONS AND THE WINDING-NUMBER 335 ourselves how many points within the simple closed curve γmap toa. The answer is given by the winding number of the image curve Γ about a. f γ Γ Figure 8.13: An analytic map is one-to-one where the winding number is unity, but two-to-one at points where the image curve winds t wice. To that this is so, we appeal to the principal of the argument a s # of zeros of ( f−a) withinγ=1 2πi/contintegraldisplay γf/prime(z) f(z)−adz, =1 2πi/contintegraldisplay Γdw w−a, =n(Γ,a), (8.142) wheren(Γ,a) is called the winding number of the image curve Γ about a. It is equal to n(Γ,a) =1 2π∆γarg (w−a), (8.143) and is the number of times the image point wencirclesaasztraverses the original curve γ. Since the number of pre-image points cannot be negative, the se winding numbers must be positive. This means that the holomorphic im age of curve winding in the anticlockwise direction is also a curve windi ng anticlockwise. For mathematicians, another important consequence of this result is that a holomorphic map is open–i.e.the holomorphic image of an open set is itself an open set. The local mapping theorem is therefore so metime called theopen mapping theorem . 8.5.2 Rouch´ e’s theorem Here we provide an effective tool for locating zeros of functi ons. 336 CHAPTER 8. COMPLEX ANALYSIS I Theorem (Rouch´ e): Let f(z)andg(z)be analytic within and on a simple closed contour γ. Suppose further that |g(z)|<|f(z)|everywhere on γ, then f(z)andf(z) +g(z)have the same number of zeros within γ. Before giving the proof, we illustrate Rouch´ e’s theorem by giving its most important corollary: the algebraic completeness of the com plex numbers, a result otherwise known as the fundamental theorem of algebra . This asserts that, ifRis sufficiently large, a polynomial P(z) =anzn+an−1zn−1+···+a0 has exactly nzeros, when counted with their multiplicity, lying within t he circle|z|=R. To prove this note that we can take Rsufficiently big that |anzn|=|an|Rn >|an−1|Rn−1+|an−2|Rn−2···+|a0| >|an−azn−1+an−2zn−2···+a0|, (8.144) on the circle|z|=R. We can therefore take f(z) =anznandg(z) = an−azn−1+an−2zn−2···+a0in Rouch´ e. Since anznhas exactly nzeros, all lying atz= 0, within|z|=R, we conclude that so does P(z). The proof of Rouch´ e is a corollary of the principle of the arg ument. We observe that # of zeros of f+g=n(Γ,0) =1 2π∆γarg (f+g) =1 2πi∆γln(f+g) =1 2πi∆γlnf+1 2πi∆γln(1 +g/f) =1 2π∆γargf+1 2π∆γarg (1 +g/f).(8.145) Now|g/f|<1 onγ, so 1 +g/fcannot circle the origin as we traverse γ. As a consequence ∆ γarg (1 +g/f) = 0. Thus the number of zeros of f+g insideγis the same as that of falone. (Naturally, they are not usually in the same places.) The geometric part of this argument is often illustrated by a dog on a lead. If the lead has length L, and the dog’s owner stays a distance R > L away from a lamp post, then the dog cannot run round the lamp po st unless the owner does the same. 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 337 g f+g ofΓ Figure 8.14: The curve Γis the image of γunder the map f+g. If|g|<|f|, then, asztraversesγ,f+gwinds about the origin the same number of times thatfdoes. Exercise 8.8 :Jacobi Theta Function. The function θ(z|τ) is defined for Im τ > 0 by the sum θ(z|τ) =∞/summationdisplay n=−∞eiπτn2e2πinz. Show thatθ(z+1|τ) =θ(z|τ), andθ(z+τ|τ) =e−iπτ−2πizθ(z|τ). Use this infor- mation and the principle of the argument to show that θ(z|τ) has exactly one zero in each unit cell of the Bravais lattice comprising the p ointsz=m+nτ; m,n∈Z. Show that these zeros are located at z= (m+ 1/2) + (n+ 1/2)τ. Exercise 8.9 : Use Rouch´ e’s theorem to find the number of roots of the equat ion z5+ 15z+ 1 = 0 lying within the circles, i) |z|= 2, ii)|z|= 3/2. 8.6 Analytic Functions and Topology 8.6.1 The Point at Infinity Some functions, f(z) = 1/zfor example, tend to a fixed limit (here 0) as z become large, independently of in which direction we set off t owards infinity. Others, such as f(z) = expz, behave quite differently depending on what direction we take as |z|becomes large. To accommodate the former type of function, and to be able to l egiti- mately write f(∞) = 0 forf(z) = 1/z, it is convenient to add “ ∞” to the set of complex numbers. Technically, what we are doing is to c onstructing 338 CHAPTER 8. COMPLEX ANALYSIS I theone-point compactification of the locally compact space C. We often portray this extended complex plane as a sphere S2(the Riemann sphere), using stereographic projection to locate infinity at the nor th pole, and 0 at the south pole. N zP S Figure 8.15: Stereographic mapping of the complex plane to the 2-Sphere. By the phrase a neighbourhood ofz, we mean an open set containing z. We use the stereographic map to define a neighbourhood of infinity as the stere- ographic image of a neighbourhood of the north pole. With thi s definition, the extended complex plane C∪{∞} becomes topologically a sphere, and in particular, becomes a compact set. If we wish to study the behaviour of a function “at infinity,” w e use the mapz/mapsto→ζ= 1/zto bring∞to the origin, and study the behaviour of the function there. Thus the polynomial f(z) =a0+a1z+···+aNzN(8.146) becomes f(ζ) =a0+a1ζ−1+···+aNζ−N, (8.147) and so has a pole of order Nat infinity. Similarly, the function f(z) =z−3has a zero of order three at infinity, and sin zhas an isolated essential singularity there. We must be a careful about defining residues at infinity. The residue is more a property of the 1-form f(z)dzthan of the function f(z) alone, and to find the residue we need to transform the dzas well asf(z). For example, if we setz= 1/ζindz/zwe have dz z=ζd/parenleftbigg1 ζ/parenrightbigg =−dζ ζ, (8.148) 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 339 so the 1-form (1 /z)dzhas a pole at z= 0 with residue 1, and has a pole with residue−1 at infinity—even though the function 1/zhas no pole there. This 1-form viewpoint is required for compatability with th e residue theorem: The integral of 1 /zaround the positively oriented unit circle is simultane- ously minus the integral of 1 /zabout the oppositely oriented unit circle, now regarded as a a positively oriented circle enclosing the poi nt at infinity. Thus iff(z) has of pole of order Nat infinity, and f(z) =···+a−2z−2+a−1z−1+a0+a1z+a2z2+···+ANzN =···+a−2ζ2+a−1ζ+a0+a1ζ−1+a2ζ−2+···+ANζ−N (8.149) near infinity, then the residue at infinity must be defined to be −a−1, and nota1as one might na¨ ıvely have thought. Once we have allowed ∞as a point in the set we map from, it is only natural to add it to the set we map to— in other words to allow ∞as a possible value for f(z). We will set f(a) =∞, if|f(z)|becomes unboundedly large asz→ain any manner. Thus, if f(z) = 1/zwe havef(0) =∞. The map w=/parenleftbiggz−z0 z−z∞/parenrightbigg/parenleftbiggz1−z∞ z1−z0/parenrightbigg (8.150) takes z0→0, z1→1, z∞→ ∞, (8.151) for example. Using this language, the M¨ obius maps w=az+b cz+d(8.152) become one-to-one maps of S2→S2. They are the only such globally con- formal one-to-one maps. When the matrix /parenleftbigg a b c d/parenrightbigg (8.153) is an element of SU(2), the resulting one–to-one map is a rigi d rotation of the Riemann sphere. Stereographic projection is thus revea led to be the geometric origin of the spinor representations of the rotat ion group. 340 CHAPTER 8. COMPLEX ANALYSIS I If an analytic function f(z) has no essential singularities anywhere on the Riemann sphere then fisrational , meaning that it can be written as f(z) =P(z)/Q(z) for some polynomials P,Q. We begin the proof of this fact by observing that f(z) can have only a finite number of poles. If, to the contrary, fhad an infinite number of poles then the compactness of S2would ensure that the poles would have a limit point somewhere. This would be a non-isolated singularity o ff, and hence an essential singularity. Now suppose we have poles at z1,z2,...,zNwith principal parts mn/summationdisplay m=1bn,m (z−zn)m. If one of the znis∞, we first use a M¨ obius map to move it to some finite point. Then F(z) =f(z)−N/summationdisplay n=1mn/summationdisplay m=1bn,m (z−zn)m(8.154) is everywhere analytic, and therefore continuous, on S2. ButS2being com- pact andF(z) being continuous implies that Fis bounded. Therefore, by Liouville’s theorem, it is a constant. Thus f(z) =N/summationdisplay n=1mn/summationdisplay m=1bn,m (z−zn)m+C, (8.155) and this is a rational function. If we made use of a M¨ obius map to move a pole at infinity, we use the inverse map to restore the origin al variables. This manoeuvre does not affect the claimed result because M¨ o bius maps take rational functions to rational functions. The mapz/mapsto→f(z) given by the rational function f(z) =P(z) Q(z)=anzn+an−1zn−1+···a0 bnzn+bn−1zn−1+···b0(8.156) wraps the Riemann sphere ntimes around the target S2. In other words, it is an-to-one map. 8.6.2 Logarithms and Branch Cuts The function y= lnzis defined to be the solution to z= expy. Unfortu- nately, since exp 2 πi= 1, the solution is not unique: if yis a solution, so is 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 341 y+ 2πi. Another way of looking at this is that if z=ρexpiθ, withρreal, theny= lnρ+iθ, and the angle θhas the same 2 πiambiguity. Now there is no such thing as a “many valued function.” By definition, a f unction is a machine into which we plug something and get a unique output. To make lnzinto a legitimate function we must select a unique θ= argzfor eachz. This can be achieved by cutting the zplane along a curve extending from the the branch point atz= 0 all the way to infinity. Exactly where we put thisbranch cut is not important; what isimportant is that it serve as an impenetrable fence preventing us from following the contin uous evolution of the function along a path that winds around the origin. Similar branch cuts serve to make fractional powers single v alued. We define the power zαfor for non-integral αby setting zα= exp{αlnz}=|z|αeiαθ, (8.157) wherez=|z|eiθ. For the square root z1/2we get z1/2=/radicalbig |z|eiθ/2, (8.158) where/radicalbig |z|represents the positive square root of|z|. We can therefore make this single-valued by a cut from 0 to ∞. To make/radicalbig (z−a)(z−b) single valued we only need to cut from atob. (Why? — think this through!). We can get away without cuts if we imagine the functions being mapsfrom some set other than the complex plane. The new set is called a Riemann surface . It consists of a number of copies of the complex plane, one fo r each possible value of our “multivalued function.” The map from t his new surface is then single-valued, because each possible value of the fu nction is the value of the function evaluated at a point on a different copy. The co pies of the complex plane are called sheets , and are connected to each other in a manner dictated by the function. The cut plane may now be thought of a s a drawing of one level of the multilayered Riemann surface. Think of an architect’s floor plan of a spiral-floored multi-story car park: If the archite ct starts drawing at one parking spot and works her way round the central core, a t some point she will find that the floor has become the ceiling of the part al ready drawn. The rest of the structure will therefore have to be plotted on the plan of the next floor up — but exactly where she draws the division betwee n one floor and the one above is rather arbitrary. The spiral car-park is a good model for the Riemann surface of the ln zfunction. See figure 8.16. 342 CHAPTER 8. COMPLEX ANALYSIS I O Figure 8.16: Part of the Riemann surface for lnz. Each time we circle the origin, we go up one level. To see what happens for a square root, follow z1/2along a curve circling the branch point singularity at z= 0. We come back to our starting point with the function having changed sign; A second trip along the sam e path would bring us back to the original value. The square root thus has o nly two sheets, and they are cross-connected as shown in figure 8.17. O Figure 8.17: Part of the Riemann surface for√z. Two copies of Care cross- connected. Circling the origin once takes you to the lower le vel. A second circuit brings you back to the upper level. In figures 8.16 and 8.17, we have shown the cross-connections being made rather abruptly along the cuts. This is not necessary —there is no singularity in the function at the cut — but it is often a convenient way to t hink about the structure of the surface. For example, the surface for/radicalbig (z−a)(z−b) also consists of two sheets. If we include the point at infinit y, this surface can be thought of as two spheres, one inside the other, and cro ss connected along the cut from atob. 8.6.3 Topology of Riemann surfaces Riemann surfaces often have interesting topology. Indeed m uch of modern algebraic topology emerged from the need to develop tools to understand multiply-connected Riemann surfaces. As we have seen, the c omplex num- bers, with the point at infinity included, have the topology o f a sphere. The 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 343 αb c a d β Figure 8.18: The 1-cycles αandβon the plane with two square-root branch cuts. The dashed part of αlies hidden on the second sheet of the Riemann surface. /radicalbig (z−a)(z−b) surface is still topologically a sphere. To see this imagin e continuously deforming the Riemann sphere by pinching it at the equator down to a narrow waist. Now squeeze the front and back of the wa ist to- gether and (imagining that the the surface can pass freely th rough itself) fold the upper half of the sphere inside the lower. The result is th e precisely the two-sheeted/radicalbig (z−a)(z−b) surface described above. The Riemann surface of the function/radicalbig (z−a)(z−b)(z−c)(z−d), which can be thought of a two spheres, one inside the other and connected along two cuts, o ne fromato band one from ctod, is, however, a torus. Think of the torus as a bicycle inner tube. Imagine using the fingers of your left hand to pinc h the front and back of the tube together and the fingers of your right hand to d o the same on the diametrically opposite part of the tube. Now fold the t ube about the pinch lines through itself so that one half of the tube is insi de the other, and connected to the outer half through two square-root cros s-connects. If you have difficulty visualizing this process, figures 8.18 and 8.19 show how the two 1-cycles, αandβ, that generate the homology group H1(T2) appear when drawn on the plane cut from atobandctod, and then when drawn on the torus. Observe, in figure 8.18, how the curves in the two-s heeted plane manage to intersect in only one point, just as they do when dra wn on the torus in figure 8.19. That the topology of the twice-cut plane is that of a torus has important consequences. This is because the elliptic integral w=I−1(z) =/integraldisplayz z0dt/radicalbig (t−a)(t−b)(t−c)(t−d)(8.159) maps the twice-cut z-plane 1-to-1 onto the torus, the latter being considered as the complex w-plane with the points wandw+nω1+mω2identified. The 344 CHAPTER 8. COMPLEX ANALYSIS I αβ Figure 8.19: The 1-cycles αandβon the torus. two numbers ω1,2are given by ω1=/contintegraldisplay αdt/radicalbig (t−a)(t−b)(t−c)(t−d), ω2=/contintegraldisplay βdt/radicalbig (t−a)(t−b)(t−c)(t−d), (8.160) and are called the periods of the elliptic function z=I(w). The map w/mapsto→ z=I(w) is a genuine function because the original zis uniquely determined byw. It is doubly periodic because I(w+nω1+mω2) =I(w), n,m∈Z. (8.161) The inverse “function” w=I−1(z) is not a genuine function of z, however, becausewincreases by ω1orω2each timezgoes around a curve deformable intoαorβ, respectively. The periods are complicated functions of a,b,c,d . If you recall our discussion of de Rham’s theorem from chapte r 4, you will see that the ωiare the results of pairing the closed holomorphic 1-form. “dw” =dz/radicalbig (z−a)(z−b)(z−c)(z−d)∈H1(T2) (8.162) with the two generators of H1(T2). The quotation marks about dware there to remind us that dwis not an exact form, i.e.it is not the exterior derivative of a single-valued function w. This cohomological interpretation of the periods of the elliptic function is the origin of the us e of the word “period” in the context of de Rham’s theorem. (See section 10 .5 for more information on elliptic functions.) More general Riemann surfaces are oriented 2-manifolds tha t can be thought of as the surfaces of doughnuts with gholes. The number gis called 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 345 1αβ β β α α1 22 33 Figure 8.20: A surfaceMof genus 3. The non-bounding 1-cycles αiandβi form a basis of H1(M). The entire surface forms the single 2-cycle that spans H2(M). thegenus of the surface. The sphere has g= 0 and the torus has g= 1. The Euler character of the Riemann surface of genus gisχ= 2(1−g). For example, figure 8.20 shows a surface of genus three. The surfa ce is in one piece, so dim H0(M) = 1. The other Betti numbers are dim H1(M) = 6 and dimH2(M) = 1, so χ=2/summationdisplay p=0(−1)pdimHp(M) = 1−6 + 1 =−4, (8.163) in agreement with χ= 2(1−3) =−4. For complicated functions, the genus may be infinite. If we have two complex variables zandwthen a polynomial relation P(z,w) = 0 defines a complex algebraic curve . Except for degenerate cases, this one (complex) dimensional curve is simultaneously a tw o (real) dimen- sional Riemann surface. With P(z,w) =z3+ 3w2z+w+ 3 = 0, (8.164) for example, we can think of z(w) being a three-sheeted function of wdefined by solving this cubic. Alternatively we can consider w(z) to be the two- sheeted function of zobtained by solving the quadratic equation w2+1 3zw+(3 +z3) 3z= 0. (8.165) In each case the branch points will be located where two or mor e roots coincide. The roots of (8.165), for example, coincide when 1−12z(3 +z3) = 0. (8.166) 346 CHAPTER 8. COMPLEX ANALYSIS I This quartic equation has four solutions, so there are four s quare-root branch points. Although constructed differently, the Riemann surf ace forw(z) and the Riemann surface for z(w) will have the same genus (in this case g= 1) because they are really are one and the same object — the algeb raic curve defined by the original polynomial equation. In order to capture all its points at infinity, we often consid er a complex algebraic curve as being a subset of CP2. To do this we make the defining equation homogeneous by introducing a third co-ordinate. F or example, for (8.164) we make P(z,w) =z3+3w2z+w+3→P(z,w,v ) =z3+3w2z+wv2+3v3.(8.167) The points where P(z,w,v ) = 0 define7aprojective curve lying in CP2. Places on this curve where the co-ordinate vis zero are the added points at infinity. Places where vis non-zero (and where we may as well set v= 1) constitute the original affine curve . A generic (non-singular) curve P(z,w) =/summationdisplay r,sarszrws= 0, (8.168) with its points at infinity included, has genus g=1 2(d−1)(d−2). (8.169) Hered= max (r+s) is the degree of the curve. This degree-genus relation is due to Pl¨ ucker. It is not, however, trivial to prove. Also not easy to prove is Riemann’s theorem of 1852 that anyfinite genus Riemann surface is the complex algebraic curve associated with some two-variable polynomial. The two assertions in the previous paragraph seem to contrad ict each other. “Any” finite genus, must surely include g= 2, but how can a genus two surface be a complex algebraic curve? There is no integer value ofdsuch that (d−1)(d−2)/2 = 2. This is where the “non-singular” caveat becomes important. An affine curve P(z,w) = 0 is said to be singular at P = (z0,w0) if all of P(z,w),∂P ∂z,∂P ∂w, 7A homogeneous polynomial P(z,w,v ) of degree ndoes not provide a map from CP2→CbecauseP(λz,λw,λv ) =λnP(z,w,v ) usually depends on λ, while the co- ordinates (λz,λw,λv ) and (z,w,v ) correspond to the same point in CP2. The zero set whereP= 0 is, however, well-defined in CP2. 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 347 vanish at P. A projective curve is singular at P ∈CP2if all of P(z,w,v ),∂P ∂z,∂P ∂w,∂P ∂v are zero there. If the curve has a singular point then then it d egenerates and ceases to be a manifold. Now Riemann’s construction does not guarantee anembedding of the surface into CP2, only an immersion . The distinction between these two concepts is that an immersed surface is all owed to self- intersect, while an embedded one is not. Being a double root o f the defining equationP(z,w) = 0, a point of self-intersection is necessarily a singular point. As an illustration of a singular curve, consider our earlier example of the curve w2= (z−a)(z−b)(z−c)(z−d) (8.170) whose Riemann surface we know to be a torus once two some point s are added at infinity, and when a,b,c,d are all distinct. The degree-genus formula applied to this degree four curve gives, however, g= 3 instead of the expected g= 1. This is because the corresponding projective curve w2v2= (z−av)(z−bv)(z−cv)(z−dv) (8.171) has a tacnode singularity at the point ( z,w,v ) = (0,1,0). Rather than investigate this rather complicated singularity at infinit y, we will consider the simpler case of what happens if we allow bto coincide with c. Whenb andcmerge, the finite point P = ( w0,z0) = (0,b) becomes a singular. Near the singularity, the equation defining our curve looks like 0 =w2−ad(z−b)2, (8.172) which is the equation of two lines, w=√ ad(z−b) andw=−√ ad(z−b), that intersect at the point ( w,z) = (0,b). To understand what is happening topologically it is first necessary to realize that a complex line is a copy of C and hence, after the point at infinity is included, is topolog ically a sphere. A pair of intersecting complex lines is therefore topologica lly a pair of spheres sharing a common point. Our degenerate curve only looks like a pair of lines near the point of intersection however. To see the larg er picture, look back at the figure of the twice-cut plane where we see that as bapproaches cwe have an αcycle of zero total length. A zero length cycle means that 348 CHAPTER 8. COMPLEX ANALYSIS I the circumference of the torus becomes zero at P, so that it lo oks like a bent sausage with its two ends sharing the common point P. Ins tead of two separate spheres, our sausage is equivalent to a single two- sphere with two points identified. PPP αβ αβ Figure 8.21: A degenerate torus is topologically the same as a sphere with two points identified. As it stands, such a set is no longer a manifold because any nei ghbourhood of P will contain bits of both ends of the sausage, and therefore cannot be given co-ordinates that make it look like a region in R2. We can, however, simply agree to delete the common point, and then plug the resulting holes in the sausage ends with two distinct points. The new set is again a m anifold, and topologically a sphere. From the viewpoint of the pair of int ersecting lines, this construction means that we stay on one line, and ignore t he other as it passes through. A similar resolution of singularities allows us to regard immersed surfaces as non-singular manifolds, and it is this sense that Riemann ’s theorem is to be understood. When nsuch self-intersection double points are deleted and replaced by pairs of distinct points The degree-genus formu la becomes g=1 2(d−1)(d−2)−n, (8.173) and this can take any integer value. 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 349 8.6.4 Conformal geometry of Riemann surfaces In this section we recall Hodge’s theory of harmonic forms fr om section 4.7.1, and see how it looks from a complex variable perspective. Thi s viewpoint reveals a relationship between Riemann surfaces and Rieman n manifolds that forms an important ingredient in string and conformal field t heory. Isothermal co-ordinates and complex structure Suppose we have a two-dimensional orientable Riemann manif oldMwith metric ds2=gijdxidxj. (8.174) In two dimensions gijhas three independent components. When we make a co-ordinate transformation we have two arbitrary function s at our disposal, and so we can use this freedom to select local co-ordinates in which only one independent component remains. The most useful choice is isothermal (also called conformal ) co-ordinates x,yin which the metric tensor is diagonal, gij=eσδij, and so ds2=eσ(dx2+dy2). (8.175) Theeσis called the scale factor orconformal factor . If we set z=x+iy andz=x−iythe metric becomes ds2=eσ(z,z)dzdz. (8.176) We can construct isothermal co-ordinates for some open neig hbourhood of any point in M. If in an overlapping isothermal co-ordinate patch the metr ic is ds2=eτ(ζ,ζ)dζdζ, (8.177) and if the co-ordinates have the same orientation, then in th e overlap region ζmust be a function only of zandζa function only of z. This is so that eτ(ζ,ζ)dζdζ=eσ(z,z)/vextendsingle/vextendsingle/vextendsingle/vextendsingledz dζ/vextendsingle/vextendsingle/vextendsingle/vextendsingle2 dζdζ (8.178) without any dζ2ordζ2terms appearing. A manifold with an atlas of complex charts whose change-of-co-ordinate formulae are holomorp hic in this way is said to be a complex manifold , and the co-ordinates endow it with a complex 350 CHAPTER 8. COMPLEX ANALYSIS I structure . The existence of a global complex structure allows to us to d e- fine the notion of meromorphic and rational functions on M. Our Riemann manifold is therefore also a Riemann surface . While any compact, orientable, two-dimensional Riemann ma nifold has a complex structure that is determined by the metric, the map ping: metric →complex structure is not one-to-one. Two metrics gij, ˜gijthat are related by a conformal scale factor gij=λ(x1,x2)˜gij (8.179) give rise to the same complex structure. Conversely, a pair o f two-dimensional Riemann manifolds having the same complex structure have me trics that are related by a scale factor. The use of isothermal co-ordinates simplifies many computat ions. Firstly, observe that gij/√g=δij, the conformal factor having cancelled. If you look back at its definition, you will see that this means that when t he Hodge “ ⋆” map acts on one-forms, the result is independent of the metri c. Ifωis a one-form ω=pdx+qdy, (8.180) then ⋆ω=−qdx+pdy. (8.181) Note that, on one-forms, ⋆⋆=−1. (8.182) Withz=x+iy,z=x−iy, we have ω=1 2(p−iq)dz+1 2(p+iq)dz. (8.183) Let us focus on the dzpart: A=1 2(p−iq)dz=1 2(p−iq)(dx+idy). (8.184) Then ⋆A=1 2(p−iq)(dy−idx) =−iA. (8.185) Similarly, with B=1 2(p+iq)dz (8.186) 8.6. ANALYTIC FUNCTIONS AND TOPOLOGY 351 we have ⋆B=iB. (8.187) Thus thedzanddzparts of the original form are separately eigenvectors of ⋆ with different eigenvalues. We use this observation to const ruct a resolution of the identity Idinto the sum of two projection operators Id=1 2(1 +i⋆) +1 2(1−i⋆), =P +P, (8.188) wherePprojects on the dzpart andPonto thedzpart of the form. The original form is harmonic if it is both closed dω= 0, and co-closed d⋆ω= 0. Thus, in two dimensions, the notion of being harmonic ( i.e.a solution of Laplace’s equation) is independent of what metr ic we are given. Ifωis a harmonic form, then ( p−iq)dzand (p+iq)dzare separately closed. Observe that ( p−iq)dzbeing closed means that ∂z(p−iq) = 0, and so p−iq is a holomorphic (and hence harmonic) function. Since both ( p−iq) anddz depend only on z, we will call ( p−iq)dza holomorphic 1-form. The complex conjugate form (p−iq)dz= (p+iq)dz (8.189) then depends only on zand is anti-holomorphic. Riemann bilinear relations As an illustration of the interplay of harmonic forms and two -dimensional topology, we derive some famous formuæ due to Riemann. These formulæ have applications in string theory and in conformal field the ory. Suppose that Mis a Riemann surface of genus g, withαi,βi,i= 1,...,g , the representative generators of H1(M) that intersect as shown in figure 8.20. By applying Hodge-de Rham to this surface, we know that we can select a set of 2gindependent, real, harmonic, 1-forms as a basis of H1(M,R). With the aid of the projector Pwe can assemble these into gholomorphic closed 1-forms ωi, together with ganti-holomorphic closed 1-forms ωi, the original 2greal forms being recovered from these as ωi+ωiand⋆(ωi+ ωi) =i(ωi−ωi). A physical interpretation of these forms is as the zand zcomponents of irrotational and incompressible fluid flows on the surface M. It is not surprising that such flows form a 2 greal dimensional, or g complex dimensional, vector space because we can independe ntly specify the 352 CHAPTER 8. COMPLEX ANALYSIS I circulation/contintegraltext v·draround each of the 2 ggenerators of H1(M). If the flow field has (covariant) components vx,vy, thenω=vzdzwherevz= (vx−ivy)/2, andω=vzdzwherevz= (vx+ivy)/2. Suppose now that aandbare closed 1-forms on M. Then, either by exploiting the powerful and general intersection-form for mula (4.77) or by cutting open the surface along the curves αi,βiand using the more direct strategy that gave us (4.79), we find that /integraldisplay Ma∧b=g/summationdisplay i=1/braceleftbigg/integraldisplay αia/integraldisplay βib−/integraldisplay βia/integraldisplay αib/bracerightbigg . (8.190) We use this formula to derive two bilinear relations associated with a closed holomorphic 1-form ω. Firstly we compute its Hodge inner-product norm /bardblω/bardbl2≡/integraldisplay Mω∧⋆ω=g/summationdisplay i=1/braceleftbigg/integraldisplay αiω/integraldisplay βi⋆ω−/integraldisplay βiω/integraldisplay αi⋆ω/bracerightbigg =ig/summationdisplay i=1/braceleftbigg/integraldisplay αiω/integraldisplay βiω−/integraldisplay βiω/integraldisplay αiω/bracerightbigg =ig/summationdisplay i=1/braceleftbig AiBi−BiAi/bracerightbig , (8.191) whereAi=/integraltext αiωandBi=/integraltext βiω. We have used the fact that ωis an anti- holomorphic 1 form and thus an eigenvector of ⋆with eigenvalue i. It follows, therefore, that if all the Aiare zero then/bardblω/bardbl= 0 and so ω= 0. LetAij=/integraltext αiωj. The determinant of the matrix Aijis non-zero: If it were zero, then there would be numbers λi, not all zero, such that 0 =Aijλj=/integraldisplay αi(ωjλj), (8.192) but, by (8.191), this implies that /bardblωjλj/bardbl= 0 and hence ωjλj= 0, contrary to the linear independence of the ωi. We can therefore solve the equations Aijλjk=δik (8.193) for the numbers λjkand use these to replace each of the ωiby the linear combination ωjλji. The new ωithen obey/integraltext αiωj=δij. From now on we suppose that this has be done. 8.7. FURTHER EXERCISES AND PROBLEMS 353 Defineτij=/integraltext βiωj. Observe that dz∧dz= 0 forces ωi∧ωj= 0, and therefore we have a second relation 0 =/integraldisplay Mωm∧ωn=g/summationdisplay i=1/braceleftbigg/integraldisplay αiωm/integraldisplay βiωn−/integraldisplay βiωm/integraldisplay αiωn/bracerightbigg =g/summationdisplay i=1{δimτin−τimδin} =τmn−τnm. (8.194) The matrix τijis therefore symmetric. A similar compuation shows that /bardblλiωi/bardbl2= 2λi(Imτij)λj (8.195) so the matrix (Im τij) is positive definite. The set of such symmetric matrices whose imaginary part is positive definite is called the Siegel upper half-plane . Not every such matrix correponds to a Riemann surface, but wh en it does it encodes all information about the shape of the Riemann manif oldMthat is left invariant under conformal rescaling. 8.7 Further Exercises and Problems Exercise 8.10 :Harmonic partners. Show that the function u= sinxcoshy+ 2cosxsinhy is harmonic. Determine the corresponding analytic functio nu+iv. Exercise 8.11 :M¨ obius Maps. The Map z/mapsto→w=az+b cz+d is called a M¨ obius transformation. These maps are importan t because they are the only one-to-one conformal maps of the Riemann sphere ont o itself. a) Show that two successive M¨ obius transformations z/prime=az+b cz+d, z/prime/prime=Az/prime+B Cz/prime+D give rise to another M¨ obius transformation, and show that t he rule for combining them is equivalent to matrix multiplication. 354 CHAPTER 8. COMPLEX ANALYSIS I b) Letz1,z2,z3,z4be complex numbers. Show that a necessary and suf- ficient condition for the four points to be concyclic is that t heir their cross-ratio {z1,z2,z3,z4}def=(z1−z4)(z3−z2) (z1−z2)(z3−z4) be real (Hint: use a well-known property of opposite angles o f a cyclic quadrilateral). Show that M¨ obius transformations leave t he cross-ratio invariant, and thus take circles into circles. Exercise 8.12 :Hyperbolic geometry . The Riemann metric for the Poincar´ e- disc model of Lobachevski’s hyperbolic plane (See exercise s??.??and 3.13) can be taken to be ds2=4|dz|2 (1−|z|2)2,|z|2<1. a) Show that the M¨ obius transformation z/mapsto→w=eiλz−a ¯az−1,|a|<1, λ∈R provides a 1-1 map of the interior of the unit disc onto itself . Show that these maps form a group. b) Show that the hyperbolic-plane metric is left invariant u nder the group of maps in part (a). Deduce that such maps are orientation-pr eserving isometries of the hyperbolic plane. c) Use the circle-preserving property of the M¨ obius maps to deduce that circles in hyperbolic geometry are represented in the Poinc ar´ e disc by Euclidean circles that lie entirely within the disc. The conformal maps of part (a) are in fact the onlyorientation preserving isometries of the hyperbolic plane. With the exception of ci rcles centered at z= 0, the center of the hyperbolic circle does not coincide wit h the center of its representative Euclidean circle. Euclidean circles that are internally tangent to the boundary of the unit disc have infinite hyperbo lic radius and their hyperbolic centers lie on the boundary of the unit disc and hence at hyperbolic infinity. They are known as horocycles. Exercise 8.13 :Rectangle to Ellipse. Consider the map w/mapsto→z= sinw. Draw a picture of the image, in the zplane, of the interior of the rectangle with cornersu=±π/2,v=±λ. (w=u+iv). Show which points correspond to the corners of the rectangle, and verify that the vertex angl es remainπ/2. At what points does the isogonal property fail? 8.7. FURTHER EXERCISES AND PROBLEMS 355 Exercise 8.14 : The part of the negative real axis where x <−1 is occupied by a conductor held at potential −V0. The positive real axis for x >+1 is similarly occupied by a conductor held at potential + V0. The conductors extend to infinity in both directions perpendicular to the x−yplane, and so the potential Vsatisfies the two-dimensional Laplace equation. a) Find the image in the ζplane of the cut zplane where the cuts run from −1 to−∞and from +1 to + ∞under the map z/mapsto→ζ= sin−1z b) Use your answer from part a) to solve the electrostatic pro blem and show that the field lines and equipotentials are conic sectio ns of the form ax2+by2= 1. Find expressions for aandbfor the both the field lines and the equipotentials and draw a labelled sketch to illustrate your results. Exercise 8.15 : Draw the image under the map z/mapsto→w=eπz/aof the infinite stripS, consisting of those points z=x+iy∈Cfor which 0 < y < a . Label enough points to show which point in the wplane corresponds to which in thezplane. Hence or otherwise show that the Dirichlet Green func tion G(x,y;x0,y0) that obeys ∇2G=δ(x−x0)δ(y−y0) inS, andG(x,y;x0,y0) = 0 for (x,y) on the boundary of S, can be written as G(x,y;x0,y0) =1 2πln|sinh(π(z−z0)/2a)|+... The dots indicate the presence of a second function, similar to the first, that you should find. Assume that ( x0,y0)∈S. Exercise 8.16 : State Laurent’s theorem for functions analytic in an annul us. Include formulae for the coefficients of the expansion. Show t hat, suitably interpreted, this theorem reduces to a form of Fourier’s the orem for functions analytic in a neighbourhood of the unit circle. Exercise 8.17 :Laurent Paradox. Show that in the annulus 1 <|z|<2 the function f(z) =1 (z−1)(2−z) has a Laurent expansion in powers of z. Find the coefficients. The part of the series with negative powers of zdoes not terminate. Does this mean that f(z) has an essential singularity at z= 0? 356 CHAPTER 8. COMPLEX ANALYSIS I Exercise 8.18 : Assuming the following series 1 sinhz=1 z−1 6z+7 16z3+..., evaluate the integral I=/contintegraldisplay |z|=11 z2sinhzdz. Now evaluate the integral I=/contintegraldisplay |z|=41 z2sinhzdz. (Hint: The zeros of sinh zlie atz=nπi.) Exercise 8.19 : State the theorem relating the difference between the numbe r of poles and zeros of f(z) in a region to the winding number of argument of f(z). Hence, or otherwise, evaluate the integral I=/contintegraldisplay C5z4+ 1 z5+z+ 1dz whereCis the circle|z|= 2. Prove, including a statement of any relevent theorem, any assertions you make about the locations of the z eros ofz5+z+1. Exercise 8.20 :Arcsine branch cuts. Letw= sin−1z. Show that w=nπ±iln{iz+/radicalbig 1−z2} with the±being selected depending on whether nis odd or even. Where would you put cuts to ensure that wis a single-valued function? Problem 8.21 :Cutting open a genus-2 surface. The Riemann surface for the function y=/radicalbig (z−a1)(z−a2)(z−a3)(z−a4)(z−a5)(z−a6) has genusg= 2. Such a surface Mis sketched in figure 8.22, where the four independent 1-cycles α1,2andβ1,2that generate H1(M) have been drawn so that they share a common vertex. a) Realize the genus-2 surface as two copies of C∪{∞} cross-connected by three square-root branch cuts. Sketch how the 1-cycles αiandβi,i= 1,2 of figure 8.22 appear when drawn on your thrice-cut plane. 8.7. FURTHER EXERCISES AND PROBLEMS 357 1234567 8 β1 β2α2α1 Figure 8.22: Concurrent 1-cycles on a genus-2 surface. 1634 5 2 7 8α1 L β1 L α1 Rβ1 Rβ2 Rα2 Lβ2 L α2 R Figure 8.23: The cut-open genus-2 surface. The superscripts L and R denot e respectively the left and right sides of each 1-cycle, viewe d from the direction of the arrow orienting the cycle. 358 CHAPTER 8. COMPLEX ANALYSIS I b) Cut the surface open along the four 1-cycles, and show that resulting surface is homeomorphic to the octagonal region appearing i n figure 8.23. c) Apply the direct method that gave us (4.79) to the octagona l region of part b). Hence show that for closed 1-forms a,b, on the surface we have /integraldisplay Ma∧b=2/summationdisplay i=1/braceleftbigg/integraldisplay αia/integraldisplay βib−/integraldisplay βia/integraldisplay αib/bracerightbigg . Chapter 9 Complex Analysis II In this chapter we will apply what we have learned of complex v ariables. The applications will range from the elementary to the sophisti cated. 9.1 Contour Integration Technology The goal of contour integration technology is to evaluate or dinary, real- variable, definite integrals. We have already met the basic t ool, the residue theorem : Theorem: Let f(z)be analytic within and on the boundary Γ =∂Dof a simply connected domain D, with the exception of finite number of points at which the function has poles. Then /contintegraldisplay Γf(z)dz=/summationdisplay poles∈D2πi(residue at pole) . 9.1.1 Tricks of the Trade The effective application of the residue theorem is somethin g of an art, but there are useful classes of integrals which we can learn to re cognize. Rational Trigonometric Expressions Integrals of the form/integraldisplay2π 0F(cosθ,sinθ)dθ (9.1) 359 360 CHAPTER 9. COMPLEX ANALYSIS II are dealt with by writing cos θ=1 2(z+z), sinθ=1 2i(z−z) and integrating around the unit circle. For example, let a,bbe real and b<a, then I=/integraldisplay2π 0dθ a+bcosθ=2 i/contintegraldisplay |z|=1dz bz2+ 2az+b=2 ib/contintegraldisplaydz (z−α)(z−β).(9.2) Sinceαβ= 1, only one pole is within the contour. This is at α= (−a+√ a2−b2)/b. (9.3) The residue is2 ib1 α−β=1 i1√ a2−b2. (9.4) Therefore, the integral is given by I=2π√ a2−b2. (9.5) These integrals are, of course, also do-able by the “ t” substitution t= tan(θ/2), whence sinθ=2t 1 +t2,cosθ=1−t2 1 +t2, dθ =2dt 1 +t2, (9.6) followed by a partial fraction decomposition. The labour is perhaps slightly less using the contour method. Rational Functions Integrals of the form/integraldisplay∞ −∞R(x)dx, (9.7) whereR(x) is a rational function of xwith the degree of the denominator exceeding the degree of the numerator by two or more, may be ev aluated by integrating around a rectangle from −Ato +A,AtoA+iB,A+iBto −A+iB, and back down to −A. Because the integrand decreases at least as fast as 1 /|z|2aszbecomes large, we see that if we let A,B→∞, the contributions from the unwanted parts of the contour become negligeable. Thus I= 2πi/parenleftBig/summationdisplay Residues of poles in upper half-plane/parenrightBig . (9.8) 9.1. CONTOUR INTEGRATION TECHNOLOGY 361 We could also use a rectangle in the lower half-plane with the result I=−2πi/parenleftBig/summationdisplay Residues of poles in lower half-plane/parenrightBig , (9.9) This must give the same answer. For example, let nbe a positive integer and consider I=/integraldisplay∞ −∞dx (1 +x2)n. (9.10) The integrand has an n-th order pole at z=±i. Suppose we close the contour in the upper half-plane. The new contour encloses the pole at z= +iand we therefore need to compute its residue. We set z−i=ζand expand 1 (1 +z2)n=1 [(i+ζ)2+ 1]n=1 (2iζ)n/parenleftbigg 1−iζ 2/parenrightbigg−n =1 (2iζ)n/parenleftBigg 1 +n/parenleftbiggiζ 2/parenrightbigg +n(n+ 1) 2!/parenleftbiggiζ 2/parenrightbigg2 +···/parenrightBigg .(9.11) The coefficient of ζ−1is 1 (2i)nn(n+ 1)···(2n−2) (n−1)!/parenleftbiggi 2/parenrightbiggn−1 =1 22n−1i(2n−2)! ((n−1)!)2. (9.12) The integral is therefore I=π 22n−2(2n−2)! ((n−1)!)2. (9.13) These integrals can also be done by partial fractions. 9.1.2 Branch-cut integrals Integrals of the form I=/integraldisplay∞ 0xα−1R(x)dx, (9.14) whereR(x) is rational, can be evaluated by integration round a slotte d circle (or “key-hole”) contour. 362 CHAPTER 9. COMPLEX ANALYSIS II y x −1 Figure 9.1: A slotted circle contour Γof outer radius Λand inner radius /epsilon1. A little more work is required to extract the answer, though. For example, consider I=/integraldisplay∞ 0xα−1 1 +xdx, 0<Reα<1. (9.15) The restrictions on the range of αare necessary for the integral to converge at its upper and lower limits. We take Γ to be a circle of radius Λ centred at z= 0, with a slot indenta- tion designed to exclude the positive real axis, which we tak e as the branch cut ofzα−1, and a small circle of radius /epsilon1about the origin. The branch of the fractional power is defined by setting zα−1= exp[(α−1)(ln|z|+iθ)], (9.16) where we will take θto be zero immediately above the real axis, and 2 π immediately below it. With this definition the residue at the pole atz=−1 iseiπ(α−1). The residue theorem therefore tells us that/contintegraldisplay Γzα−1 1 +zdz= 2πieπi(α−1). (9.17) The integral decomposes as /contintegraldisplay Γzα−1 1 +zdz=/contintegraldisplay |z|=Λzα−1 1 +zdz+ (1−e2πi(α−1))/integraldisplayΛ /epsilon1xα−1 1 +xdx−/contintegraldisplay |z|=/epsilon1zα−1 1 +zdz. (9.18) 9.1. CONTOUR INTEGRATION TECHNOLOGY 363 As we send Λ off to infinity we can ignore the “1” in the denominat or com- pared to the z, and so estimate /vextendsingle/vextendsingle/vextendsingle/vextendsingle/contintegraldisplay |z|=Λzα−1 1 +zdz/vextendsingle/vextendsingle/vextendsingle/vextendsingle→/vextendsingle/vextendsingle/vextendsingle/vextendsingle/contintegraldisplay |z|=Λzα−2dz/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤2πΛ×ΛRe(α)−2. (9.19) This tends to zero provided that Re α<1. Similarly, provided 0 <Reα, the integral around the small circle about the origin tends to ze ro with/epsilon1. Thus −eπiα2πi=/parenleftbig 1−e2πi(α−1)/parenrightbig I. (9.20) We conclude that I=2πi (eπiα−e−πiα)=π sinπα. (9.21) Exercise 9.1 : Using the slotted circle contour, show that I=/integraldisplay∞ 0xp−1 1 +x2dx=π 2sin(πp/2)=π 2cosec (πp/2),0<p< 2. Exercise 9.2 : Integrate za−1/(z−1) around a contour Γ 1consisting of a semi- circle in the upper half plane together with the real axis ind ented atz= 0 andz= 1 xy 1 Figure 9.2: The contour Γ1. to get 0 =/contintegraldisplay Γza−1 z−1dz=P/integraldisplay∞ 0xa−1 x−1dx−iπ+ (cosπa+isinπa)/integraldisplay∞ 0xa−1 x+ 1dx. 364 CHAPTER 9. COMPLEX ANALYSIS II As usual, the symbol Pin front of the integral sign denotes a principal part integral, meaning that we must omit an infinitesimal segment of the contour symmetrically disposed about the pole at z= 1. The term−iπcomes from integrating around the small semicircle about this point. W e get−1/2 of the residue because we have only a half circle, and that traverse d in the “wrong” direction. Warning : this fractional residue result is only true when we indent to avoid a simple pole —i.e.one that is of order one. Now take real and imaginary parts and deduce that /integraldisplay∞ 0xa−1 1 +xdx=π sinπα,0<Rea<1, and P/integraldisplay∞ 0xa−1 1−xdx=πcotπa, 0<Rea<1. 9.1.3 Jordan’s Lemma We often need to evaluate Fourier integrals I(k) =/integraldisplay∞ −∞eikxR(x)dx (9.22) withR(x) a rational function. For example, the Green function for th e operator−∂2 x+m2is given by G(x) =/integraldisplay∞ −∞dk 2πeikx k2+m2. (9.23) Supposex∈Randx >0. Then, in contrast to the analogous integral without the exponential function, we have no flexibility in c losing the contour in the upper or lower half-plane. The function eikxgrows without limit as we head south in the lower half-plane, but decays rapidly in t he upper half- plane. This means that we may close the contour without chang ing the value of the integral by adding a large upper-half-plane semicirc le. 9.1. CONTOUR INTEGRATION TECHNOLOGY 365 Rk im −im Figure 9.3: Closing the contour in the upper half-plane. The modified contour encloses a pole at k=im, and this has residue i/(2m)e−mx. Thus G(x) =1 2me−mx, x> 0. (9.24) Forx<0, the situation is reversed, and we must close in the lower ha lf-plane. The residue of the pole at k=−imis−i/(2m)emx, but the minus sign is cancelled because the contour goes the “wrong way” (clockwi se). Thus G(x) =1 2me+mx, x< 0. (9.25) We can combine the two results as G(x) =1 2me−m|x|. (9.26) The formal proof that the added semicircles make no contribu tion to the integral when their radius becomes large is known as Jordan’s Lemma : Lemma: Let Γbe a semicircle, centred at the origin, and of radius R. Sup- pose i) thatf(z)is meromorphic in the upper half-plane; ii) thatf(z)tends uniformly to zero as |z|→∞ for0<argz<π; iii) the number λis real and positive. Then /integraldisplay Γeiλzf(z)dz→0,asR→∞. (9.27) 366 CHAPTER 9. COMPLEX ANALYSIS II To establish this, we assume that Ris large enough that |f|< /epsilon1on the contour, and make a simple estimate /vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplay Γeiλzf(z)dz/vextendsingle/vextendsingle/vextendsingle/vextendsingle<2R/epsilon1/integraldisplayπ/2 0e−λRsinθdθ <2R/epsilon1/integraldisplayπ/2 0e−2λRθ/πdθ =π/epsilon1 λ(1−e−λR)<π/epsilon1 λ. (9.28) In the second inequality we have used the fact that (sin θ)/θ≥2/πfor angles in the range 0 <θ<π/ 2. Since/epsilon1can be made as small as we like, the lemma follows. Example : Evaluate I(α) =/integraldisplay∞ −∞sin(αx) xdx. (9.29) We have I(α) = Im/braceleftbigg/integraldisplay∞ −∞expiαz zdz/bracerightbigg . (9.30) If we takeα>0, we can close in the upper half-plane, but our contour must exclude the pole at z= 0. Therefore 0 =/integraldisplay |z|=Rexpiαz zdz−/integraldisplay |z|=/epsilon1expiαz zdz+/integraldisplay−/epsilon1 −Rexpiαx xdx+/integraldisplayR /epsilon1expiαx xdx. (9.31) AsR→∞, we can ignore the big semicircle, the rest, after letting /epsilon1→0, gives 0 =−iπ+P/integraldisplay∞ −∞eiαx xdx. (9.32) Again, the symbol Pdenotes a principal part integral. The −iπcomes from the small semicircle. We get −1/2 the residue because we have only a half circle, and that traversed in the “wrong” direction. (Remem ber that this fractional residue result is only true when we indent to avoi d asimple pole — i.eone that is of order one.) Reading off the real and imaginary parts, we conclude that /integraldisplay∞ −∞sinαx xdx=π, P/integraldisplay∞ −∞cosαx xdx= 0, α> 0. (9.33) 9.1. CONTOUR INTEGRATION TECHNOLOGY 367 No “P” is needed in the sine integral, as the integrand is finite at x= 0. If we relax the condition that α>0 and take into account that sine is an odd function of its argument, we have /integraldisplay∞ −∞sinαx xdx=πsgnα. (9.34) This identity is called Dirichlet’s discontinuous integral . We can interpret Dirichlet’s integral as giving the Fourier transform of the principal part distribution P(1/x) as P/integraldisplay∞ −∞eiωx xdx=iπsgnω. (9.35) This will be of use later in the chapter. Example : xy Figure 9.4: Quadrant contour. We will evaluate the integral /contintegraldisplay Ceizza−1dz (9.36) about the first-quadrant contour shown above. Observe that w hen 0<a< 1 neither the large nor the small arc makes a contribution, and that there are no poles. Hence, we deduce that 0 =/integraldisplay∞ 0eixxa−1dx−i/integraldisplay∞ 0e−yya−1e(a−1)π 2idy,0<a< 1. (9.37) 368 CHAPTER 9. COMPLEX ANALYSIS II Taking real and imaginary parts, we find /integraldisplay∞ 0xa−1cosxdx = Γ(a) cos/parenleftBigπ 2a/parenrightBig ,0<a< 1, /integraldisplay∞ 0xa−1sinxdx = Γ(a) sin/parenleftBigπ 2a/parenrightBig ,0<a< 1, (9.38) where Γ(a) =/integraldisplay∞ 0ya−1e−ydy (9.39) is the Euler Gamma function. Example: Fresnel integrals . Integrals of the form C(t) =/integraldisplayt 0cos(πx2/2)dx, (9.40) S(t) =/integraldisplayt 0sin(πx2/2)dx, (9.41) occur in the theory of diffraction and are called Fresnel integrals after Au- gustin Fresnel. They are naturally combined as C(t) +iS(t) =/integraldisplayt 0eiπx2/2dx. (9.42) The limit as t→∞ exists and is finite. Even though the integrand does not tend to zero at infinity, its rapid oscillation for large xis just sufficient to ensure convergence.1 Astvaries, the complex function C(t)+iS(t) traces out the Cornu Spiral , named after Marie Alfred Cornu, a 19th century French optica l physicist. 1We can exhibit this convergence by setting x2=sand then integrating by parts to get /integraldisplayt 0eiπx2/2dx=1 2/integraldisplay1 0eiπs/2ds s1/2+/bracketleftbiggeiπs/2 πis1/2/bracketrightbiggt2 1+1 2πi/integraldisplayt2 1eiπs/2ds s3/2. The right hand side is now manifestly convergent as t→∞. 9.1. CONTOUR INTEGRATION TECHNOLOGY 369 -0.75 -0.5 -0.25 0.25 0.5 0.75 -0.6-0.4-0.20.20.40.6 Figure 9.5: The Cornu spiral C(t)+iS(t)fortin the range−8<t< 8. The spiral in the first quadrant corresponds to positive values o ft. We can evaluate the limiting value C(∞) +iS(∞) =/integraldisplay∞ 0eiπx2/2dx (9.43) by deforming the contour off the real axis and onto a line of len gthLrunning into the first quadrant at 45◦, this being the direction of most rapid decrease of the integrand. Ly x Figure 9.6: Fresnel contour. A circular arc returns the contour to the axis whence it conti nues to∞, but an estimate similar to that in Jordan’s lemma shows that the a rc and the 370 CHAPTER 9. COMPLEX ANALYSIS II subsequent segment on the real axis make a negligeable contr ibution when L is large. To evaluate the integral on the radial line we set z=eiπ/4s, and so /integraldisplayeiπ/4∞ 0eiπz2/2dz=eiπ/4/integraldisplay∞ 0e−πs2/2ds=1√ 2eiπ/4=1 2(1 +i).(9.44) Figure 9.5 shows how C(t) +iS(t) orbits the limiting point 0 .5 + 0.5iand slowly spirals in towards it. Taking real and imaginary part s we have /integraldisplay∞ 0cos/parenleftbiggπx2 2/parenrightbigg dx=/integraldisplay∞ 0sin/parenleftbiggπx2 2/parenrightbigg dx=1 2. (9.45) 9.2 The Schwarz Reflection Principle Theorem (Schwarz): Let f(z)be analytic in a domain Dwhere∂Dincludes a segment of the real axis. Assume that f(z)is real when zis real. Then there is a unique analytic continuation of finto the region D(the mirror image ofDin the real axis) given by g(z) =  f(z), z∈D, f(z), z∈D, either, z∈R.(9.46) xy D D Figure 9.7: The domain Dand its mirror image D. The proof invokes Morera’s theorem to show analyticity, and then appeals to the uniqueness of analytic continuations. Begin by looki ng at a closed 9.2. THE SCHWARZ REFLECTION PRINCIPLE 371 contour lying only in D:/contintegraldisplay Cf(z)dz, (9.47) whereC={η(t)}is the image of C={η(t)}⊂Dunder reflection in the real axis. We can rewrite this as /contintegraldisplay Cf(z)dz=/contintegraldisplay f(η)d¯η dtdt=/contintegraldisplay f(η)dη dtdt=/contintegraldisplay Cf(η)dz= 0. (9.48) At the last step we have used Cauchy and the analyticity of finD. Morera’s theorem therefore confirms that g(z) is analytic in D. By breaking a general contour up into parts in Dand parts in D, we can similarly show that g(z) is analytic in D∪D. The important corollary is that if f(z) is analytic, and real on some segment of the real axis, but has a cut along some other part of the real axis, thenf(x+i/epsilon1) =f(x−i/epsilon1) as we go over the cut. The discontinuity disc fis therefore 2Im f(x+i/epsilon1). Supposef(z) is real on the negative real axis, and goes to zero as |z|→∞ , then applying Cauchy to the contour Γ depicted in figure 9.8. y xζ Figure 9.8: The contour Γfor the dispersion relation. . we find f(ζ) =1 π/integraldisplay∞ 0Imf(x+i/epsilon1) x−ζdx, (9.49) 372 CHAPTER 9. COMPLEX ANALYSIS II forζwithin the contour. This is an example of a dispersion relation . The name comes from the prototypical application of this techno logy to optical dispersion, i.e.the variation of the refractive index with frequency. Iff(z) does not tend to zero at infinity then we cannot ignore the con - tribution to Cauchy’s formula from the large circle. We can, however, still write f(ζ) =1 2πi/contintegraldisplay Γf(z) z−ζdz, (9.50) and f(b) =1 2πi/contintegraldisplay Γf(z) z−bdz, (9.51) for some convenient point bwithin the contour. We then subtract to get f(ζ) =f(b) +(ζ−b) 2πi/integraldisplay Γf(z) (z−b)(z−ζ)dz. (9.52) Because of the extra power of zdownstairs in the integrand, we only need f to be bounded at infinity for the contribution of the large cir cle to tend to zero. If this is the case, we have f(ζ) =f(b) +(ζ−b) π/integraldisplay∞ 0Imf(x+i/epsilon1) (x−b)(x−ζ)dx. (9.53) This is called a once-subtracted dispersion relation. The dispersion relations derived above apply when ζlies within the con- tour. In physics applications we often need f(ζ) forζreal and positive. What happens as ζapproaches the axis, and we attempt to divide by zero in such an integral, is summarized by the Plemelj formulæ : Iff(ζ) is defined by f(ζ) =1 π/integraldisplay Γρ(z) z−ζdz, (9.54) where Γ has a segment lying on the real axis, then, if xlies in this segment, 1 2(f(x+i/epsilon1)−f(x−i/epsilon1)) =iρ(x) 1 2(f(x+i/epsilon1) +f(x−i/epsilon1)) =P π/integraldisplay Γρ(x/prime) x/prime−xdx/prime. (9.55) As always, the “ P” means that we are to delete an infinitesimal segment of the contour lying symmetrically about the pole. 9.2. THE SCHWARZ REFLECTION PRINCIPLE 373 + = 2− = Figure 9.9: Origin of the Plemelj formulae. The Plemelj formulæ hold under relatively mild conditions o n the function ρ(x). We won’t try to give a general proof, but in the case that ρis analytic the result is easy to understand: we can push the contour out o f the way and letζ→xon the real axis from either above or below. In that case the drawing above shows how the the sum of these two limits giv es the the principal-part integral and how their difference gives an in tegral round a small circle, and hence the residue ρ(x). The Plemelj equations usually appear in physics papers as th e “i/epsilon1” cabala 1 x/prime−x±i/epsilon1=P/parenleftbigg1 x/prime−x/parenrightbigg ∓iπδ(x/prime−x). (9.56) A limit/epsilon1→0 is always to be understood in this formula. Im fRe f x’−xx’−x Figure 9.10: Sketch of the real and imaginary parts of f(x/prime) = 1/(x/prime−x−i/epsilon1). We can also appreciate the origin of the i/epsilon1rule by examining the following identity: 1 x/prime−(x±i/epsilon1)=x−x/prime (x/prime−x)2+/epsilon12±i/epsilon1 (x/prime−x)2+/epsilon12. (9.57) 374 CHAPTER 9. COMPLEX ANALYSIS II The first term is a symmetrically cut-off version of 1 /(x/prime−x) and provides the principal-part integral. The second term sharpens and t ends to the delta function±iπδ(x/prime−x) as/epsilon1→0. Exercise 9.3 : The Legendre function of the second kind Qn(z) may be defined for positive integer nby the integral Qn(z) =1 2/integraldisplay1 −1(1−t2)n 2n(z−t)n+1dt, z /∈[−1,1]. Show that for x∈[−1,1] we have Qn(x+i/epsilon1)−Qn(x−i/epsilon1) =−iπPn(x), wherePn(x) is the Legendre Polynomial. Deduce Neumann ’s formula Qn(z) =1 2/integraldisplay1 −1Pn(t) z−tdt, z /∈[−1,1]. 9.2.1 Kramers-Kronig Relations Causality is the usual source of analyticity in physical app lications. If G(t) is a response function φresponse (t) =/integraldisplay∞ −∞G(t−t/prime)fcause(t/prime)dt/prime(9.58) then for no effect to anticipate its cause we must have G(t) = 0 fort <0. The Fourier transform G(ω) =/integraldisplay∞ −∞eiωtG(t)dt, (9.59) is then automatically analytic everywhere in the upper half plane. Suppose, for example, we look at a forced, damped, harmonic oscillato r whose dis- placementx(t) obeys ¨x+ 2γ˙x+ (Ω2+γ2)x=F(t), (9.60) where the friction coefficient γis positive. As we saw earlier, the solution is of the form x(t) =/integraldisplay∞ −∞G(t,t/prime)F(t/prime)dt/prime, 9.2. THE SCHWARZ REFLECTION PRINCIPLE 375 where the Green function G(t,t/prime) = 0 ift<t/prime. In this case G(t,t/prime) =/braceleftBigg Ω−1e−γ(t−t/prime)sin Ω(t−t/prime)t>t/prime 0, t<t/prime(9.61) and so x(t) =1 Ω/integraldisplayt −∞e−γ(t−t/prime)sin Ω(t−t/prime)F(t/prime)dt/prime. (9.62) Because the integral extends only from 0 to + ∞, the Fourier transform of G(t,0), ˜G(ω)≡1 Ω/integraldisplay∞ 0eiωte−γtsin Ωtdt, (9.63) is nicely convergent when Im ω>0, as evidenced by ˜G(ω) =−1 (ω+iγ)2−Ω2(9.64) having no singularities in the upper half-plane.2 Another example of such a causal function is provided by the c omplex, frequency-dependent, refractive index of a material n(ω). This is defined so that a travelling wave takes the form ϕ(x,t) =ein(ω)k·x−iωt. (9.65) We can decompose ninto its real and imaginary parts n(ω) =nR(ω) +inI(ω) =nR(ω) +i 2|k|γ(ω) (9.66) whereγis the extinction coefficient, defined so that the intensity fa lls off asI∝exp(−γn·x), where n=k/|k|is the direction of propapagation. A non-zeroγcan arise from either energy absorption or scattering out of the forward direction 2If a pole in a response function manages to sneak into the uppe r half plane, then the system will be unstable to exponentially growing oscill ations. This may happen, for example, when we design an electronic circuit containing a f eedback loop. Such poles, and the resultant instabilities, can be detected by applying th e principle of the argument from the last chapter. This method leads to the Nyquist stability criterion. 376 CHAPTER 9. COMPLEX ANALYSIS II Being a causal response, the refractive index extends to a fu nction ana- lytic in the upper half plane and n(ω) for realωis the boundary value n(ω)physical = lim /epsilon1→0n(ω+i/epsilon1) (9.67) of this analytic function. Because a real ( E=E∗) incident wave must give rise to a real wave in the material, and because the wave must d ecay in the direction in which it is propagating, we have the reality con ditions γ(−ω+i/epsilon1) =−γ(ω+i/epsilon1), nR(−ω+i/epsilon1) = +nR(ω+i/epsilon1) (9.68) withγpositive for positive frequency. Many materials have a frequency range |ω|<|ωmin|whereγ= 0, so the material is transparent. For any such material n(ω) obeys the Schwarz reflection principle and so there is an analytic continuatio n into the lower half-plane. At frequencies ωwhere the material is not perfectly transparent, the refractive index has an imaginary part even when ωis real. By Schwarz, n must be discontinuous across the real axis at these frequenc ies:n(ω+i/epsilon1) = nR+inI/negationslash=n(ω−i/epsilon1) =nR−inI. These discontinuities of 2 inIusually correspond to branch cuts. No substance is able to respond to infinitely high frequency d isturbances, son→1 as|ω|→∞ , and we can apply our dispersion relation technology to the function n−1. We will need the contour shown below, which has cuts for both positive and negative frequencies. Im Reω ω ωmin −ωmin Figure 9.11: Contour for the n−1dispersion relation. 9.2. THE SCHWARZ REFLECTION PRINCIPLE 377 By applying the dispersion-relation strategy, we find n(ω) = 1 +1 π/integraldisplayωmin −∞nI(ω/prime) ω/prime−ωdω/prime+1 π/integraldisplay∞ ωminnI(ω/prime) ω/prime−ωdω/prime(9.69) forωwithin the contour. Using Plemelj we can now take ωonto the real axis to get nR(ω) = 1 +P π/integraldisplayωmin −∞nI(ω/prime) ω/prime−ωdω/prime+P π/integraldisplay∞ ωminnI(ω/prime) ω/prime−ωdω/prime = 1 +P π/integraldisplay∞ ω2 minnI(ω/prime) ω/prime2−ω2dω/prime2, = 1 +c πP/integraldisplay∞ ωminγ(ω/prime) ω/prime2−ω2dω/prime. (9.70) In the second line we have used the anti-symmetry of nI(ω) to combine the positive and negative frequency range integrals. In the las t line we have used the relation ω/k=cto make connection with the way this equation is written in R. G. Newton’s authoritative Scattering Theory of Waves and Particles . This relation, between the real and absorptive parts of the r efractive index, is called a Kramers-Kronig dispersion relation, after the original authors.3 Ifn→1 fast enough that ω2(n−1)→0 as|ω|→∞ , we can take the f in the dispersion relation to be ω2(n−1) and deduce that nR= 1 +c πP/integraldisplay∞ ω2 min/parenleftbiggω/prime2 ω2/parenrightbiggγ(ω/prime) ω/prime2−ω2dω/prime, (9.71) another popular form of Kramers-Kronig. This second relati on implies the first, but not vice-versa , because the second demands more restrictive be- havior forn(ω). Similar equations can be derived for other causal functions . A quantity closely related to the refractive index is the frequency-de pendent dielectric “constant” /epsilon1(ω) =/epsilon11+i/epsilon12. (9.72) Again/epsilon1→1 as|ω|→∞ , and, proceeding as before, we deduce that /epsilon11(ω) = 1 +P π/integraldisplay∞ ω2 min/epsilon12(ω/prime) ω/prime2−ω2dω/prime2. (9.73) 3H. A. Kramers, Nature ,117(1926) 775; R. de L. Kronig, J. Opt. Soc. Am. 12(1926) 547 378 CHAPTER 9. COMPLEX ANALYSIS II 9.2.2 Hilbert transforms Suppose that f(x) is the boundary value on the real axis of a function every- where analytic in the upper half-plane, and suppose further thatf(z)→0 as|z|→∞ there. Then we have f(z) =1 2πi/integraldisplay∞ −∞f(x) x−zdx (9.74) forzin the upper half-plane. This is because may close the contou r with an upper semicircle without changing the value of the integral . For the same reason the integral must give zero when zis taken in the lower half-plane. Using the Plemelj formulæ we deduce that on the real axis, f(x) =P πi/integraldisplay∞ −∞f(x/prime) x/prime−xdx/prime. (9.75) We can use this strategy to derive the Kramers-Kronig relati ons even if nI never vanishes, and so we cannot use the Schwarz reflection pr inciple. The relation (9.75) suggests the definition of the Hilbert transform ,Hψ, of a function ψ(x), as (Hψ)(x) =P π/integraldisplay∞ −∞ψ(x/prime) x−x/primedx/prime. (9.76) Note the interchange of x,x/primein the denominator of (9.76) when compared with (9.75). This switch is to make the Hilbert transform int o a convolution integral. Equation (9.75) shows that a function that is the b oundary value of a function analytic and tending to zero at infinity in the uppe r half-plane is automatically an eigenvector of Hwith eigenvalue−i. Similarly a function that is the boundary value of a function analytic and tending to zero at infinity in the lower half-plane will be an eigenvector with e igenvalue + i. (A function analytic in the entire complex plane and tending to zero at infinity must vanish identically by Liouville’s theorem.) Returning now to our original f, which had eigenvalue −i, and decom- posing it as f(x) =fR(x) +ifI(x) we find that (9.75) becomes fI(x) = (HfR)(x), fR(x) =−(HfI)(x). (9.77) 9.2. THE SCHWARZ REFLECTION PRINCIPLE 379 Conversely, if we are given a real function u(x) and setv(x) = (Hu)(x), then, under some mild restrictions on u(that it lie in some Lp(R),p>1, for example, in which case v(x) is also in Lp(R).) the function f(z) =1 2πi/integraldisplay∞ −∞u(x) +iv(x) x−zdx (9.78) will be analytic in the upper half plane, tend to zero at infini ty there, and haveu(x) +iv(x) as its boundary value as zapproaches the real axis from above. The last line of (9.77) therefore shows that we may rec overu(x) from v(x) asu(x) =−(Hv)(x). The Hilbert transform H:Lp(R)→Lp(R) is therefore invertible, and its inverse is given by H−1=−H. (Note that the Hilbert transform of a constant is zero, but the Lp(R) condition excludes constants from the domain of H, and so this fact does not conflict with invertibility.) Hilbert transforms are useful in signal processing. Given a real signal XR(t) we can take its Hilbert transform so as to find the correspond ing imaginary part, XI(t), which serves to make the sum Z(t) =XR(t) +iXI(t) =A(t)eiφ(t)(9.79) analytic in the upper half-plane. This complex function is t heanalytic sig- nal.4The real quantity A(t) is then known as the instantaneous amplitude , orenvelope , whileφ(t) is the instantaneous phase and ωIF(t) =˙φ(t) (9.80) is called the instantaneous frequency (IF). These quantities are used, for example, in narrow band FM radio, in NMR, in geophysics, and i n image processing. Exercise 9.4 : Let/tildewidef(ω) =/integraltext∞ −∞eiωtf(t)dtdenote the Fourier transform of f(t). Use the formula (9.35) for the Fourier transform of P(1/t), combined with the convolution theorem for Fourier transforms, to show that th e Fourier transform of the Hilbert transform of f(t) is /tildewide(Hf)(ω) =isgn(ω)/tildewidef(ω). Deduce that the analytic signal is derived from the original real signal by suppressing all positive frequency components (those prop ortional to e−iωt withω>0) and multiplying the remaining negative-frequency ampli tudes by two. 4D. Gabor, J. Inst. Elec. Eng. (Part 3) ,93(1946) 429-457. 380 CHAPTER 9. COMPLEX ANALYSIS II Exercise 9.5 : Suppose that ϕ1(x) andϕ2(x) are real functions with finite L2(R) norms. a) Use the Fourier transform result from the previous exerci se to show that /angbracketleftϕ1,ϕ2/angbracketright=/angbracketleftHϕ1,Hϕ2/angbracketright. Thus,His a unitary transformation from L2(R)→L2(R). b) Use the fact that H2=−Ito deduce that /angbracketleftHϕ1,ϕ2/angbracketright=−/angbracketleftϕ1,Hϕ2/angbracketright and soH†=−H. c) Conclude from part b) that /integraldisplay∞ −∞ϕ1(x)/parenleftbigg P/integraldisplay∞ −∞ϕ2(y) x−ydy/parenrightbigg dx=/integraldisplay∞ −∞ϕ2(y)/parenleftbigg P/integraldisplay∞ −∞ϕ1(x) x−ydx/parenrightbigg dy, i.e., forL2(R), functions, it is legitimate to interchange the order of “ P” integration with ordinary integration. d) By replacing ϕ1(x) by a constant, and ϕ2(x) by the Hilbert transform of a function fwith/integraltext fdx/negationslash= 0, show that it is not always safe to interchange the order of “ P” integration with ordinary integration Exercise 9.6 : Suppose that are given real functions u1(x) andu2(x) and sub- stitute their Hilbert transforms v1=Hu1,v2=Hu2into (9.78) to construct analytic functions f1(z) andf2(z). Then the product f1(z)f2(z) =F(z) has boundary value FR(x) +iFI(x) = (u1u2−v1v2) +i(u1v2+u2v1). By assuming that F(z) satisfies the conditions for (9.77) to be applicable to this boundary value, deduce that H((Hu1)u2) +H((Hu2)u1)−(Hu1)(Hu2) =−u1u2. ⋆ This result5sometimes appears in the physics literature6in the guise of the distributional identity P x−yP y−z+P y−zP z−x+P z−xP x−y=−π2δ(x−y)δ(x−z), ⋆⋆ 5F. G. Tricomi, Quart. J. Math. (Oxford) , (2)2, (1951) 199. 6For example, in R. Jackiw, A. Strominger, Phys. Lett. 99B(1981) 133. 9.3. PARTIAL-FRACTION AND PRODUCT EXPANSIONS 381 whereP/(x−y) denotes the principal-part distribution P/parenleftBig 1/(x−y)/parenrightBig . This attractively symmetric form conceals the fact that a specifi c order of inte- gration is to be understood. As the next exercise shows, were we to freely re-arrange the integration order we could use the identity 1 x−y1 y−z+1 y−z1 z−x+1 z−x1 x−y= 0 to wrongly conclude that the right-hand side is zero. Exercise 9.7 : Show that the identity ⋆from exercise 9.6 can be written as /integraldisplay∞ −∞/parenleftbigg/integraldisplay∞ −∞ϕ1(y)ϕ2(z) (z−y)(y−x)dz/parenrightbigg dy=/integraldisplay∞ −∞/parenleftbigg/integraldisplay∞ −∞ϕ1(y)ϕ2(z) (z−y)(y−x)dy/parenrightbigg dz−π2ϕ1(x)ϕ2(x), principal-part integrals being understood where necessar y. This is a special case of a more general change-of-integration-order formul a /integraldisplay∞ −∞/parenleftbigg/integraldisplay∞ −∞f(x,y,z) (z−y)(y−x)dz/parenrightbigg dy=/integraldisplay∞ −∞/parenleftbigg/integraldisplay∞ −∞f(x,y,z) (z−y)(y−x)dy/parenrightbigg dz−π2f(x,x,x ), which is due to G. H. Hardy (1908). Show that Hardy’s formula i s equivalent to the distributional identity ⋆⋆. Exercise 9.8 : Use the licit interchange of “ P” integration with ordinary inte- gration to show that /integraldisplay∞ −∞ϕ(x)/parenleftbigg P/integraldisplay∞ −∞ϕ(y) x−ydy/parenrightbigg2 dx=π2 3/integraldisplay∞ −∞ϕ3(x)dx. Exercise 9.9 : Letf(z) be analytic within the unit circle, and let u(θ) and v(θ) be the boundary values of its real and imaginary parts, resp ectively, at z=eiθ. Use Plemelj to show that u(θ) =−1 2πP/integraldisplay2π 0v(θ/prime)cot/parenleftbiggθ−θ/prime 2/parenrightbigg dθ/prime+1 2π/integraldisplay2π 0u(θ/prime)dθ/prime, v(θ) =1 2πP/integraldisplay2π 0u(θ/prime)cot/parenleftbiggθ−θ/prime 2/parenrightbigg dθ/prime+1 2π/integraldisplay2π 0v(θ/prime)dθ/prime. 9.3 Partial-Fraction and Product Expansions In this section we will study other useful representations o f functions which devolve from their analyticity properties. 382 CHAPTER 9. COMPLEX ANALYSIS II 9.3.1 Mittag-Leffler Partial-Fraction Expansion Letf(z) be a meromorphic function with poles (perhaps infinitely ma ny) atz=zj, (j= 1,2,3,...), where|z1|<|z2|< .... Let Γnbe a contour enclosing the first npoles. Suppose further (for ease of description) that the poles are simple and have residue rn. Then, for zinside Γn, we have 1 2πi/contintegraldisplay Γnf(z/prime) z/prime−zdz/prime=f(z) +n/summationdisplay j=1rj zj−z. (9.81) We often want to to apply this formula to trigonometric funct ions whose periodicity means that they do not tend to zero at infinity. We therefore employ the same subtraction strategy that we used for dispersion relations. We subtract f(z)−f(0) =z 2πi/contintegraldisplay Γnf(z/prime) z/prime(z/prime−z)dz/prime+n/summationdisplay j=1rj/parenleftbigg1 z−zj+1 zj/parenrightbigg .(9.82) If we now assume that f(z) is uniformly bounded on the Γ n— this meaning that|f(z)|< Aon Γn, with the same constant Aworking for all n— then the integral tends to zero as nbecomes large, yielding the partial fraction, orMittag-Leffler , decomposition f(z) =f(0) +∞/summationdisplay j=1rj/parenleftbigg1 z−zj+1 zj/parenrightbigg (9.83) Example 1) : Look at cosec z. The residues of 1 /(sinz) at its poles at z=nπ arern= (−1)n. We can take the Γ nto be squares with corners ( n+1/2)(±1± i)π. A bit of effort shows that cosec is uniformly bounded on them. To use the formula as given, we first need subtract the pole at z= 0, then cosecz−1 z=∞/summationdisplay n=−∞/prime (−1)n/parenleftbigg1 z−nπ+1 nπ/parenrightbigg . (9.84) The prime on the summation symbol indicates that we are omit t hen= 0 term. The positive and negative nseries converge separately, so we can add them, and write the more compact expression cosecz=1 z+ 2z∞/summationdisplay 1(−1)n1 z2−n2π2. (9.85) 9.3. PARTIAL-FRACTION AND PRODUCT EXPANSIONS 383 Example 2) : A similar method gives cotz=1 z+∞/summationdisplay n=−∞/prime/parenleftbigg1 z−nπ+1 nπ/parenrightbigg . (9.86) We can pair terms together to writen this as cotz=1 z+∞/summationdisplay n=1/parenleftbigg1 z−nπ+1 z+nπ/parenrightbigg , =1 z+∞/summationdisplay n=12z z2−n2π2(9.87) or cotz= lim N→∞N/summationdisplay n=−N1 z−nπ. (9.88) In the last formula it is important that the upper and lower li mits of summa- tion be the same. Neither the sum over positive nnor the sum over negative nconverges separately. By taking asymmetric upper and lower limits we could therefore obtain any desired number as the limit of the sum. Exercise 9.10 : Use Mittag-Leffler to show that cosec2z=∞/summationdisplay n=−∞1 (z+nπ)2. Now use this infinite series to give a one-line proof of the tri gonometric identity N−1/summationdisplay m=0cosec2/parenleftBig z+mπ N/parenrightBig =N2cosec2(Nz). (Is there a comparably easy elementary derivation of this finite sum?) Take a limit to conclude that N−1/summationdisplay m=1cosec2/parenleftBigmπ N/parenrightBig =1 3(N2−1). Exercise 9.11 : From the partial fraction expansion for cot z, deduce that d dzln[(sinz)/z] =d dz∞/summationdisplay n=1ln(z2−n2π2). 384 CHAPTER 9. COMPLEX ANALYSIS II Integrate this along a suitable path from z= 0, and so conclude that that sinz=z∞/productdisplay n=1/parenleftbigg 1−z2 n2π2/parenrightbigg . Exercise 9.12 : By differentiating the partial fraction expansion for cot z, show that, forkan integer≥1, and Imz >0, we have ∞/summationdisplay n=−∞1 (z+n)k+1=(−2πi)k+1 k!∞/summationdisplay n=1nke2πinz. This is called Lipshitz’ formula . Exercise 9.13 : The Bernoulli numbers are defined by x ex−1= 1 +B1x+∞/summationdisplay k=1B2kx2k (2k)!. The first few are B1=−1/2,B2= 1/6,B4=−1/30. Except for B1, theBn are zero for nodd. Show that xcotx=ix+2ix e2ix−1= 1−∞/summationdisplay k=1(−1)k+1B2k22kx2k (2k)!. By expanding 1 /(x2−n2π2) as a power series in xand comparing coefficients, deduce that, for positive integer k, ∞/summationdisplay n=11 n2k= (−1)k+1π2k22k−1 (2k)!B2k. Exercise 9.14 :Euler-Maclaurin sum formula . Use the formal expansion D eD−1=/summationdisplay kBkDk k!= 1−1 2D+1 6D2 2!−1 30D4 4!+···, withDinterpreted as d/dx, to obtain (−f/prime(x)−f/prime(x+1)−f/prime(x+2)+···) =f(x)−1 2f/prime(x)+1 6f/prime/prime(x) 2!−1 30f(4) 4!+···. By integrating this from atob≡a+m, motivate the Euler-Maclaurin formula m−1/summationdisplay k=0f(a+k) =/integraldisplayb af(x)dx+1 2(f(a)−f(b)) +∞/summationdisplay k=1B2k (2k)!(f(2k−1)(a)−f(2k−1)(b)). This “derivation,” while suggestive, is only heuristic. It gives no insight into whether the series converges (it usually does not) or what th e error might be if we truncate after a finite number of terms. 9.3. PARTIAL-FRACTION AND PRODUCT EXPANSIONS 385 9.3.2 Infinite Product Expansions We can play a variant of the Mittag-Leffler game with suitable e ntire func- tionsg(z) and derive for them a representation as an infinite product. Sup- pose thatg(z) has simple zeros at zi. Then (lng)/prime=g/prime(z)/g(z) is meromor- phic with poles at zi, all with unit residues. Assuming that it satisfies the uniform boundedness condition, we now use Mittag Leffler to wr ite d dzlng(z) =g/prime(z) g(z)/vextendsingle/vextendsingle/vextendsingle/vextendsingle z=0+∞/summationdisplay j=1/parenleftbigg1 z−zj+1 zj/parenrightbigg . (9.89) Integrating up we have lng(z) = lng(0) +cz+∞/summationdisplay j=1/parenleftbigg ln(1−z/zj) +z zj/parenrightbigg , (9.90) wherec=g/prime(0)/g(0). We now re-exponentiate to get g(z) =g(0)ecz∞/productdisplay j=1/parenleftbigg 1−z zj/parenrightbigg ez/zj. (9.91) Example : Letg(z) = sinz/z, theng(0) = 1, while the constant c, which is the logarithmic derivative of gatz= 0, is zero, and sinz z=∞/productdisplay n=1/parenleftBig 1−z nπ/parenrightBig ez/nπ/parenleftBig 1 +z nπ/parenrightBig e−z/nπ. (9.92) Thus sinz=z∞/productdisplay n=1/parenleftbigg 1−z2 n2π2/parenrightbigg . (9.93) Convergence of Infinite Products We have derived several infinite problem formulæ without dis cussing the issue of their convergence. For products of terms of the form (1+ an) with positive anwe can reduce the question of convergence to that of/summationtext∞ n=1an. To see why this is so, let pN=N/productdisplay n=1(1 +an), an>0. (9.94) 386 CHAPTER 9. COMPLEX ANALYSIS II Then we have the inequalities 1 +N/summationdisplay n=1an<pN<exp/braceleftBiggN/summationdisplay n=1an/bracerightBigg . (9.95) The infinite sum and product therefore converge or diverge to gether. If P=∞/productdisplay n=1(1 +|an|), (9.96) converges, we say that p=∞/productdisplay n=1(1 +an), (9.97) converges absolutely. As with infinite sums, absolute conve rgence implies convergence, but not vice-versa. Unlike infinite sums, howe ver, an infinite product containing negative ancan diverge to zero. If (1 +an)>0 then/producttext(1 +an) converges if/summationtextln(1 +an) does, and we will say that/producttext(1 +an) diverges to zero if/summationtextln(1 +an) diverges to−∞. Exercise 9.15 : Show that N/productdisplay n=1/parenleftbigg 1 +1 n/parenrightbigg =N+ 1, N/productdisplay n=2/parenleftbigg 1−1 n/parenrightbigg =1 N. From these deduce that∞/productdisplay n=2/parenleftbigg 1−1 n2/parenrightbigg =1 2. Exercise 9.16 : For|z|<1, show that ∞/productdisplay n=0/parenleftbig 1 +z2n/parenrightbig =1 1−z. (Hint: think binary) Exercise 9.17 : For|z|<1, show that ∞/productdisplay n=1(1 +zn) =∞/productdisplay n=11 1−z2n−1. (Hint: 1−x2n= (1−xn)(1 +xn).) 9.4. WIENER-HOPF EQUATIONS II 387 9.4 Wiener-Hopf Equations II The theory of Hilbert transforms has shown us some the conseq uences of functions being analytic in the upper or lower half-plane. A nother applica- tion of these ideas is to Wiener-Hopf equations . Although we have discussed Wiener-Hopf integral equations in chapter ??, it is only now that we pos- sess the tools to appreciate the general theory. We begin, ho wever, with the slightly simpler Wiener-Hopf sumequations, which are their discrete analogue. Here, analyticity in the upper or lower half-plan e is replaced by analyticity within or without the unit circle. 9.4.1 Wiener-Hopf Sum Equations Consider the infinite system of equations yn=∞/summationdisplay m=−∞an−mxm,−∞<n<∞ (9.98) where we are given the ynand are seeking the xn. If thean,ynare the Fourier coefficients of smooth complex-valued func- tions A(θ) =∞/summationdisplay n=−∞aneinθ, Y(θ) =∞/summationdisplay n=−∞yneinθ, (9.99) then the systems of equations is, in principle at least, easy to solve. We introduce the function X(θ) =∞/summationdisplay n=−∞xneinθ, (9.100) and (9.98) becomes Y(θ) =A(θ)X(θ). (9.101) From this, the desired xnmay be read off as the Fourier expansion coefficients ofY(θ)/A(θ). We see that A(θ) must be nowhere zero or else the operator A represented by the infinite matrix an−mwill not be invertible. This technique 388 CHAPTER 9. COMPLEX ANALYSIS II is a discrete version of the Fourier transform method for sol ving the integral equation y(s) =/integraldisplay∞ −∞A(s−t)x(t)dt,−∞<s<∞. (9.102) The connection with complex analysis is made by regarding A(θ),X(θ),Y(θ) as being functions on the unit circle in the zplane. If they are smooth enough we can extend their definition to an annulus about the unit cir cle, so that A(z) =∞/summationdisplay n=−∞anzn, X(z) =∞/summationdisplay n=−∞xnzn, Y(z) =∞/summationdisplay n=−∞ynzn. (9.103) Thexnmay now be read off as the Laurent expansion coefficients of Y(z)/A(z). The discrete analogue of the Wiener-Hopf integral equation y(s) =/integraldisplay∞ 0A(s−t)x(t)dt,0≤s<∞ (9.104) is the Wiener-Hopf sum equation yn=∞/summationdisplay m=0an−mxm,0≤n<∞. (9.105) This requires a more sophisticated approach. If you look bac k at our earlier discussion of Wiener-Hopf integral equations in chapter ??, you will see that the trick for solving them is to extend the definition y(s) to negative s(anal- ogously, the ynto negative n) and find these values at the same time as we findx(s) for positive s(analogously, the xnfor positive n.) We proceed by introducing the same functions A(z),X(z),Y(z) as before, but now keep careful track of whether their power-series exp ansions contain positive or negative powers of z. In doing so, we will discover that the Fredholm alternative governing the existence and uniquene ss of the solutions will depend on the winding number N=n(Γ,0) where Γ is the image of the unit circle under the map z/mapsto→A(z) — in other words, on how many times A(z) wraps around the origin as zgoes once round the unit circle. 9.4. WIENER-HOPF EQUATIONS II 389 Suppose that A(z) is smooth enough that it is analytic in an annulus including the unit circle, and that we can factorize A(z) so that A(z) =λq+(z)zN[q−(z)]−1, (9.106) where q+(z) = 1 +∞/summationdisplay n=1q+ nzn, q−(z) = 1 +∞/summationdisplay n=1q− −nz−n. (9.107) Here we demand that q+(z) be analytic and non-zero for |z|<1 +/epsilon1, and thatq−(z) be analytic and non-zero for |1/z|<1 +/epsilon1. These no pole, no zero, conditions ensure, viathe principle of the argument, that the winding numbers of q±(z) about the origin are zero, and so all the winding of A(z) is accounted for by the N-fold winding of the zNfactor. The non-zero condition also ensures that the reciprocals [ q±(z)]−1have same class of expansions ( i.e. in positive or negative powers of zonly) as the direct functions. We now introduce the notation [ F(z)]+and [F(z)]−, meaning that we expandF(z) as a Laurent series and retain only the positive powers of z (includingz0), or only the negative powers (starting from z−1), respectively. ThusF(z) = [F(z)]++[F(z)]−. We will write Y±(z) = [Y(z)]±, and similarly forX(z). We can therefore rewrite (9.105) in the form λzNq+(z)X+= [Y+(z) +Y−(z)]q−(z). (9.108) IfN≥0, and we break this equation into its positive and negative p owers, we find [Y+q−]+=λzNq+(z)X+, [Y+q−]−=−Y−q−(z). (9.109) From the first of these equations we can read off the desired xnas the positive- power Laurent coefficients of X+(z) = [Y+q−]+(λzNq+(z))−1. (9.110) As a byproduct, the second alows us to find the coefficient y−nofY−(z). Observe that there is a condition on Y+for this to work: the power series 390 CHAPTER 9. COMPLEX ANALYSIS II expansion of λzNq+(z)X+starts with zN, and so for a solution to exist the firstNterms of (Y+q−)+as a power series in zmust be zero. The given vectorynmust therefore satisfy Nconsistency conditions. A formal way of expressing this constraint begins by observing that it mean s that the range of the operator Arepresented by the matrix an−mfalls short, by Ndimensions, of the being the entire space of possible yn. This is exactly the situation that the notion of a “cokernel” is intended to capture. Recall tha t ifA:V→V, then Coker A=V/ImA. We therefore have dim [CokerA] =N. WhenN <0, on the other hand, we have [Y+(z)q−(z)]+= [λz−|N|q+(z)X+(z)]+ [Y+(z)q−(z)]−=−Y−(z)q−(z) + [λz−|N|q+(z)X+(z)]−.(9.111) Here the last term in the second equation contains no more tha nNterms. Be- cause of the z−|N|, we can add any to X+any multiple of Z+(x) =zn[q+(z)]−1 forn= 0,...,N−1,and still have a solution. Thus the solution is not unique. Instead, we have dim [Ker ( A)] =|N|. We have therefore shown that Index (A)def= dim (Ker A)−dim (Coker A) =−N This connection between a topological quantity – in the pres ent case the winding number — and the difference in dimension of the kernel and cokernel is an example of an index theorem. We now need to show that we can indeed factorize A(z) in the desired manner. When A(z) is a rational function, the factorization is straightfor- ward: if A(z) =C/producttext n(z−an)/producttext m(z−bm), (9.112) we simply take q+(z) =/producttext |an|>0(1−z/an)/producttext |bm|>0(1−z/bm), (9.113) where the products are over the linear factors correspondin g to poles and zeros outside the unit circle, and q−(z) =/producttext |bm|<0(1−bm/z)/producttext |an|<0(1−an/z), (9.114) 9.4. WIENER-HOPF EQUATIONS II 391 containing the linear factors corresponding to poles and ze ros inside the unit circle. The constant λand the power zNin equation (9.106) are the factors that we have extracted from the right-hand sides of (9.113) a nd (9.114), respectively, in order to leave 1’s as the first term in each li near factor. More generally, we take the logarithm of z−NA(z) =λq+(z)(q−(z))−1(9.115) to get ln[z−NA(z)] = ln[λq+(z)]−ln[q−(z)], (9.116) where we desire ln[ λq+(z)] to be the boundary value of a function analytic within the unit circle, and ln[ q−(z)] the boundary value of function analytic outside the unit circle and with q−(z) tending to unity as |z|→∞ . The factor ofz−Nin the logarithm serves to undo the winding of the argument ofA(z), and results in a single-valued logarithm on the unit circl e. Plemelj now shows that Q(z) =1 2πi/contintegraldisplay |z|=1ln[ζ−NA(ζ)] ζ−zdζ (9.117) provides us with the desired factorization. This function Q(z) is everywhere analytic except for a branch cut along the unit circle, and it s branches, Q+ within and Q−without the circle, differ by ln[ z−NA(z)]. We therefore have λq+(z) =eQ+(z), q−(z) =eQ−(z). (9.118) The expression for Qas an integral shows that Q(z)∼const./z as|z| goes to infinity and so guarantees that q−(z) has the desired limit of unity there. The task of finding this factorization is known as the scalar Riemann- Hilbert problem . In effect, we are decomposing the infinite matrix A= ............ ···a0a1a2··· ···a−1a0a1··· ···a−2a−1a0··· ............ (9.119) 392 CHAPTER 9. COMPLEX ANALYSIS II into the product of an upper triangular matrix U=λ ............ ···1q+ 1q+ 2··· ···0 1q+ 1··· ···0 0 1···............ , (9.120) a lower triangular matrix L, where L−1= ............ ··· 1 0 0··· ···q− −11 0··· ···q− −2q− −11··· ............ , (9.121) has 1’s on the diagonal, and a matrix ΛNwhich which is zero everywhere except for a line of 1’s located Nsteps above the main diagonal. The set of triangular matrices with unit diagonal form a group, so th e inversion required to obtain Lresults in a matrix of the same form. The resulting Birkhoff factorization A=LΛNU, (9.122) is an infinite-dimensional extension of the Gauss-Bruhat (o r generalized LU) decomposition of a matrix. The finite-dimensional Gauss-Br uhat decompo- sition provides a factorization of a matrix A∈GL(n) as A=LΠU, (9.123) whereLis a lower triangular matrix with 1’s on the diagonal, Uis an upper triangular matrix with no zero’s on the diagonal, and Πis a permutation matrix, i.e.a matrix that permutes the basis vectors by having one entry o f 1 in each row and in each column, and all other entries zero. Ou r present ΛN is playing the role of such a matrix. The matrix Πis uniquely determined byA. The LandUmatrices become unique if Lis chosen so that ΠTLΠ is also lower triangular. 9.4. WIENER-HOPF EQUATIONS II 393 9.4.2 Wiener-Hopf Integral Equations We now carry over our insights from the simpler sum equations to Weiner- Hopf integral equations /integraldisplay∞ 0K(x−y)φ(y)dy=f(x), x> 0, (9.124) by imagining replacing the unit circle by a circle of radius R, and then taking R→∞ in such a way that the sums go over to integrals. In this way man y features are retained: the problem is still solved by factor izing the Fourier transform /tildewideK(k) =/integraldisplay∞ −∞K(x)eikxdx (9.125) of the kernel, and there remains an index theorem dim (KerK)−dim (Coker K) =−N, (9.126) butNnow counts the winding of the phase of /tildewideK(k) askranges over the real axis: N=1 2πarg/tildewideK/vextendsingle/vextendsingle/vextendsinglek=+∞ k=−∞. (9.127) One restriction arises though: we will require Kto be of the form K(x−y) =δ(x−y) +g(x−y) (9.128) for some continuous function g(x). Our discussion is therefore being re- stricted to Wiener-Hopf Integral equations of the second kind . The restriction comes about about because we will seek to obt ain a fac- torization of/tildewideKas τ(κ)/tildewideK(k) = exp{Q+(k)−Q−(k)}=q+(k)(q−(k))−1(9.129) whereq+(k)≡exp{Q+(k)}is analytic and non-zero in the upper half k-plane andq−(k)≡exp{Q−(k)}analytic and non-zero in the lower half-plane. The factorτ(κ) is a phase such as τ(k) =/parenleftbiggk+i k−i/parenrightbiggN , (9.130) 394 CHAPTER 9. COMPLEX ANALYSIS II which winds−Ntimes and serves serves to undo the + Nphase winding in /tildewideK. TheQ±(k) will be the boundary values from above and below the real axis, respectively, of Q(k) =1 2πi/integraldisplay∞ −∞ln[τ(κ)/tildewideK(κ)] κ−kdκ (9.131) The convergence of this infinite integral requires that ln[ τ(κ)/tildewideK(k)] go to zero at infinity, or, in other words, lim k→∞/tildewideK(k) = 1. (9.132) This, in turn, requires that the original K(x) contain a delta function. Example : We will solve the problem φ(x)−λ/integraldisplay∞ 0e−|x−y|−α(x−y)φ(y)dy=f(x), x> 0. (9.133) We require that 0 < α < 1. The upper bound on αis necessary for the integral kernel to be bounded. We will also assume for simpli city thatλ < 1/2. Following the same strategy as in the sum case, we extend th e integral equation to the entire range of xby writing φ(x)−λ/integraldisplay∞ 0e−|x−y|−α(x−y)φ(y)dy=f(x) +g(x), (9.134) wheref(x) is nonzero only for x >0 andg(x) is non-zero only for x <0. The Fourier transform of this equation is /parenleftbigg(k+iα)2+a2 (k+iα)2+ 1/parenrightbigg /tildewideφ+(k) =/tildewidef+(k) +/tildewideg−(k), (9.135) wherea2= 1−2λand the±subscripts are to remind us that /tildewideφ(k) and/tildewidef(k) are analytic in the upper half-plane, and /tildewideg(k) in the lower. We will use the notationH+for the space of functions analytic in the upper half plane, a nd H−for functions analytic in the lower half plane, and so /tildewideφ+(k),/tildewidef(+k)∈H+,/tildewideg−(k)∈H− (9.136) We can factorize /tildewideK(k) =(k+iα)2+a2 (k+iα)2+ 1=[k+i(α−a)] [k+i(α−1)][k+i(α+a)] [k+i(α+ 1)](9.137) 9.4. WIENER-HOPF EQUATIONS II 395 Now suppose that ais small enough that α±a >0 and so the numerator has two zeros in the lower half plane, and the numerator a one z ero in each of the upper and lower half-planes. The change of phase in /tildewideK(k) as we go from minus to plus infinity is therefore −2π, and so the index is N=−1. We should therefore multiply /tildewideKby τ(k) =/parenleftbiggk+i k−i/parenrightbigg−1 (9.138) before seeking to break it into its q±factors. We can however equally well take τ(k) =/parenleftbiggk+i(α−1) k+i(α−a)/parenrightbigg (9.139) as this also undoes the winding and allows us to factorize wit h q−(k) = 1, q +(k) =/parenleftbiggk+i(α+a) k+i(α+ 1)/parenrightbigg . (9.140) The resultant equation analagous to (9.108) is therefore /parenleftbiggk+i(α+a) k+i(α+ 1)/parenrightbigg /tildewideφ+=/parenleftbiggk+i(α−1) k+i(α−a)/parenrightbigg /tildewidef++/parenleftbiggk+i(α−1) k+i(α−a)/parenrightbigg /tildewideg− q+/tildewideφ+= (τq−)/tildewidef+ +τq−/tildewideg− (9.141) The second line of this equation shows the interpretation of the first line in terms of the objects in the general theory. The left hand side is inH+— i.e.analytic in the upper half-plane. The first term on the right i s also in H+. (We are lucky. More generally it would have to be decomposed into its H±parts.) If it were not for the τ(κ), the last term would be in H−, but it has a potential pole at k=−i(α−a). We therefore remove this pole by substracting a term −β k+i(α−a) (an element of H+) from each side of the equation before projecting onto the H±parts. After projecting, we find that H+:/parenleftbiggk+i(α+a) k+i(α+ 1)/parenrightbigg /tildewideφ+−/parenleftbiggk+i(α−1) k+i(α−a)/parenrightbigg /tildewidef+−β k+i(α−a)= 0, H−:/parenleftbiggk+i(α−1) k+i(α−a)/parenrightbigg /tildewideg−−β k+i(α−a)= 0. (9.142) 396 CHAPTER 9. COMPLEX ANALYSIS II We solve for/tildewideφ+(k) and/tildewideg−(k) /tildewideφ+(k) =/parenleftbigg(k+iα)2+ 1 (k+iα)2+a2/parenrightbigg /tildewidef−−β/parenleftbiggk+i(α+ 1) (k+iα)2+a2/parenrightbigg /tildewideg−(k) =β k+i(α−1). (9.143) Observeg−(k) is always in H−because its only singularity is in the upper half-plane for any β. The constant βis therefore arbitrary. Finally, we invert the Fourier transform, using F/parenleftbig θ(x)e−αxsinhax/parenrightbig =−a (k+iα)2+a2,(α±a)>0, (9.144) to find that φ(x) =f(x)−2λ a/integraldisplayx 0e−α(x−y)sinha(x−y)f(y)dy +β/prime/braceleftbig (a−1)e−(α+a)x+ (a+ 1)e−(α−a)x/bracerightbig ,(9.145) whereβ/prime(proportional to β) is an arbitrary constant. By takingαin the range−1< α < 0 with (α±a)<0, we make index to beN= +1. We will then find there is condition on f(x) for the solution to exist. This condition is, of course, that f(x) be orthogonal to the solution φ0(x) =/braceleftbig (a−1)e−(α+a)x+ (a+ 1)e−(α−a)x/bracerightbig (9.146) of the homogenous adjoint problem, this being the f(x) = 0 case of the α>0 problem that we have just solved. 9.5 Further Exercises and Problems Exercise 9.18 :Contour Integration : Use the calculus of residues to evaluate the following integrals: I1=/integraldisplay2π 0dθ (a+bcosθ)2,0<b<a. I2=/integraldisplay2π 0cos23θ 1−2acos 2θ+a2dθ, 0<a< 1. I3=/integraldisplay∞ 0xα (1 +x2)2dx,−1<α< 2. 9.5. FURTHER EXERCISES AND PROBLEMS 397 These are not meant to be easy! You will have to dig for the resi dues. Answers: I1=2πa (a2−b2)3/2, I2=π(a3+ 1) a2−1=π(1−a+a2) a−1, I3=π(1−α) 4cos(πα/2). Exercise 9.19 : By considering the integral of f(z) = ln(1−e2iz) = ln(−2ieizsinz) around the indented rectangle iY +iYπ 0 π Figure 9.12: Indented rectangle. with vertices 0, π,π+iY,iY, and letting Ybecome large, evaluate the integral I=/integraldisplayπ 0ln(sinx)dx. Explain how the fact that /epsilon1ln/epsilon1→0 as/epsilon1→0 allows us to ignore contributions from the small indentations. You should also provide justifi cation for any other discarded contributions. Take care to make consistent choi ces of the branch of the logarithm, especially if expanding ln( −2ieixsinx) =ix+ ln 2 + ln(sin x) + ln(−i). The value of Iis a real number. 398 CHAPTER 9. COMPLEX ANALYSIS II Exercise 9.20 : By integrating a suitable function around the quadrant con - taining the point z0=eiπ/4, evaluate the integral I(α) =/integraldisplay∞ 0xα−1 1 +x4dx 0<α< 4. (It should only be necessary to consider the residue at z0.) Exercise 9.21 : In section ??we considered the causal Green function for the damped harmonic oscillator G(t) =/braceleftbigg1 Ωe−γtsin(Ωt), t> 0, 0, t< 0, and showed that its Fourier transform /integraldisplay∞ −∞eiωtG(t)dt=1 Ω2−(ω+iγ)2, (9.147) had no singularities in the upper half-plane. Use Jordan’s l emma to compute the inverse Fourier transform 1 2π/integraldisplay∞ −∞e−iωt Ω2−(ω+iγ)2dω, and verify that it reproduces G(t). Problem 9.22 :Jordan’s Lemma and one-dimensional scattering theory . In problem ??.??we considered the one-dimensional scattering problem solu tions ψk(x) =/braceleftbigg eikx+rL(k)e−ikx, x∈L, tL(k)eikx, x∈R,k>0. =/braceleftbigg tR(k)eikx, x∈L, eikx+rR(k)e−ikx, x∈R.k<0. and claimed that the bound-state contributions to the compl eteness relation were given in terms of the reflection and transmission coeffici ents as /summationdisplay boundψ∗ n(x)ψn(x/prime) =−/integraldisplay∞ −∞dk 2πrL(k)e−ik(x+x/prime), x,x/prime∈L, =−/integraldisplay∞ −∞dk 2πtL(k)e−ik(x−x/prime), x∈L, x/prime∈R, =−/integraldisplay∞ −∞dk 2πtR(k)e−ik(x−x/prime), x∈R, x/prime∈L, =−/integraldisplay∞ −∞dk 2πrR(k)e−ik(x+x/prime), x,x/prime∈R. 9.5. FURTHER EXERCISES AND PROBLEMS 399 The eigenfunctions ψ(+) k(x) =/braceleftbigg eikx+rL(k)e−ikx, x∈L, tL(k)eikx, x∈R, and ψ(−) k(x) =/braceleftbigg tR(k)eikx, x∈L, eikx+rR(k)e−ikx, x∈R. are initially refined for kreal and positive ( ψ(+) k) or forkreal and negative (ψ(−) k), but they separately have analytic continuations to all of k∈C. The reflection and transmission coefficients rL,R(k) andtL,R(k) are also analytic functions of k, and obey rL,R(k) =r∗ L,R(−k∗),tL,R(k) =t∗ L,R(−k∗). a) By inspecting the formulæ for ψ(+) k(x), show that the bound states ψn(x), withEn=−κ2 n, are proportional to ψ(+) k(x) evaluated at points k=iκn on the positive imaginary axis at which rL(k) andtL(k) simultaneously have poles. Similarly show that these same bound states are p roportional toψ(−) k(x) evaluated at points −iκnon the negative imaginary axis at whichrR(k) andtR(k) have poles. (All these functions ψ(±) k(x),rR,L(k), tR,L(k), may have branch points and other singularities in the half -plane on the opposite side of the real axis from the bound-state pol es.) b) Use Jordan’s lemma to evaluate the Fourier transforms giv en above in terms of the position and residues of the bound-state poles. Confirm that your answers are of the form /summationdisplay nA∗ n[sgn(x)]e−κn|x|An[sgn(x/prime)]e−κn|x/prime|, as you would expect for the bound-state contribution to the c ompleteness relation. Exercise 9.23 :Lattice Matsubara sums : Show that sums over the N-th roots of−1 can be written as an integral 1 N/summationdisplay ωN+1=0f(ω) =1 2πi/integraldisplay Cdz zzN zN+ 1f(z), whereCconsists of a pair of oppositely oriented concentric circle s. The annu- lus formed by the circles should include all the roots of unit y, but exclude all singularites of f. Use this trick to show that, for Neven, 1 NN−1/summationdisplay n=0sinhE sinh2E+ sin2(2n+1)π N=1 coshEtanhNE 2. 400 CHAPTER 9. COMPLEX ANALYSIS II Take theN→∞ limit in some suitable manner, and hence show that ∞/summationdisplay n=−∞a a2+ [(2n+ 1)π]2=1 2tanha 2. (Hint: If you are careless, you will end up differing by a facto r of two from this last formula. There are tworegions in the finite sum that tend to the infinite sum in the large Nlimit.) Problem 9.24 : If we define χ(h) =eαxφ(x), andF(x) =eαxf(x), then the Wiener-Hopf equation φ(x)−λ/integraldisplay∞ 0e−|x−y|−α(x−y)φ(y)dy=f(x), x> 0. becomes χ(x)−λ/integraldisplay∞ 0e−|x−y|χ(y)dy=F(x), x> 0, all mention of αhaving disappeared! Why then does our answer, worked out in such detail, in section 9.4.2 depend on the parameter α? Show that if α small enough that α+ais positive and α−ais negative, then φ(x) really is independent of α. (Hint: What tacit assumptions about function spaces does our use of Fourier transforms entail? How does the inverse Fo urier transform of [(k+iα)2+a2]−1vary withα?) Chapter 10 Special Functions II In this chapter we will apply complex analytic methods so as t o obtain a wider view of some of the special functions of mathematical p hysics than can be obtained on the real axis. The standard text in this field re mains the venerable Course of Modern Analysis of E. T. Whittaker and G. N. Watson. 10.1 The Gamma Function We begin with Euler’s “Gamma Function” Γ( z). You probably have some acquaintance with this creature. The usual definition is Γ(z) =/integraldisplay∞ 0tz−1e−tdt,Rez >0,(definition A) . (10.1) An integration by parts, based on d dt/parenleftbig tze−t/parenrightbig =ztz−1e−t−tze−t, (10.2) shows that/bracketleftbig tze−t/bracketrightbig∞ 0=z/integraldisplay∞ 0tz−1e−tdt−/integraldisplay∞ 0tze−tdt. (10.3) The integrated out part vanishes at both limits, provided th e real part of z is greater than zero. Thus Γ(z+ 1) =zΓ(z). (10.4) 401 402 CHAPTER 10. SPECIAL FUNCTIONS II Since Γ(1) = 1, we deduce that Γ(n) = (n−1)!, n= 1,2,3,···. (10.5) We can use the recurrence relation to extend the definition of Γ(z) to the left half plane, where the real part of zis negative. Choosing an integer nsuch that the real part of z+nis positive, we write Γ(z) =Γ(z+n) z(z+ 1)···(z+n−1). (10.6) We see that Γ( z) has poles at zero, and at the negative integers. The residue of the pole at z=−nis (−1)n/n!. We can also view the analytic continuation as an example of Ta ylor series subtraction. Let us recall how this works. Suppose that −1<Rex <0. Then, from d dt(txe−t) =xtx−1e−t−txe−t(10.7) we have/bracketleftbig txe−t/bracketrightbig∞ /epsilon1=x/integraldisplay∞ /epsilon1dttx−1e−t−/integraldisplay∞ /epsilon1dttxe−t. (10.8) Here we have cut off the integral at the lower limit so as to avoi d the di- vergence near t= 0. Evaluating the left-hand side and dividing by xwe find −1 x/epsilon1x=/integraldisplay∞ /epsilon1dttx−1e−t−1 x/integraldisplay∞ /epsilon1dttxe−t. (10.9) Since, for this range of x, −1 x/epsilon1x=/integraldisplay∞ /epsilon1dttx−1, (10.10) we can rewrite (10.9) as 1 x/integraldisplay∞ /epsilon1dttxe−t=/integraldisplay∞ /epsilon1dttx−1/parenleftbig e−t−1/parenrightbig . (10.11) The integral on the right-hand side of this last expression i s convergent as /epsilon1→0, so we may safely take the limit and find 1 xΓ(x+ 1) =/integraldisplay∞ 0dttx−1/parenleftbig e−t−1/parenrightbig . (10.12) 10.1. THE GAMMA FUNCTION 403 Since the left-hand side is equal to Γ( x), we have shown that Γ(x) =/integraldisplay∞ 0dttx−1/parenleftbig e−t−1/parenrightbig ,−1<Rex<0. (10.13) Similarly, if−2<Rex<−1, we can show that Γ(x) =/integraldisplay∞ 0dttx−1/parenleftbig e−t−1 +t/parenrightbig . (10.14) Thus the analytic continuation of the original integral is g iven by a new integral in which we have subtracted exactly as many terms fr om the Taylor expansion of e−tas are needed to just make the integral convergent. Other useful identities, usually proved by elementary real -variable meth- ods, include Euler’s “Beta function” identity, B(a,b)def=Γ(a)Γ(b) Γ(a+b)=/integraldisplay1 0(1−t)a−1tb−1dt (10.15) (which, as the Veneziano formula , was the original inspiration for string theory) and Γ(z)Γ(1−z) =πcosecπz. (10.16) The proofs of both formulæ begin in the same way: set t=y2,x2, so that Γ(a)Γ(b) = 4/integraldisplay∞ 0y2a−1e−y2dy/integraldisplay∞ 0x2b−1e−x2dx = 4/integraldisplay∞ 0/integraldisplay∞ 0e−(x2+y2)x2b−1y2a−1dxdy = 2/integraldisplay∞ 0e−r2(r2)a+b−1d(r2)/integraldisplayπ/2 0sin2a−1θcos2b−1θdθ. We have appealed to Fubini’s theorem twice: once to turn a pro duct of integrals into a double integral, and once (after setting x=rcosθ,y= rsinθ) to turn the double integral back into a product of decoupled integrals. In the second factor of the third line we can now change variab les tot= sin2θ and obtain the Beta function identity. If, on the other hand, we puta= 1−z, b=zwe have Γ(z)Γ(1−z) = 2/integraldisplay∞ 0e−r2d(r2)/integraldisplayπ/2 0cot2z−1θdθ= 2/integraldisplayπ/2 0cot2z−1θdθ. (10.17) 404 CHAPTER 10. SPECIAL FUNCTIONS II Now set cot θ=ζ. The last integral then becomes (see exercise 9.1): 2/integraldisplay∞ 0ζ2z−1 ζ2+ 1dζ=πcosecπz, 0<z < 1. (10.18) Although this integral has a restriction on the range of z, the result (10.16) can be analytically continued to so as to hold for all z. If we put z= 1/2 we find that (Γ(1 /2))2=π. The positive square root is the correct one, and Γ(1/2) =√π. (10.19) The integral in definition A is only convergent for Re z >0. A more powerful definition, involving an integral which converges for allz, is 1 Γ(z)=1 2πi/integraldisplay Cet tzdt.(definition B) (10.20) CRe(t) Im(t) Figure 10.1: Definition “B” contour for Γ(z). HereCis a contour originating at z=−∞−i/epsilon1, below the negative real axis (on which a cut serves to make t−zsingle valued) rounding the origin, and then heading back to z=−∞+i/epsilon1— this time staying above the cut. We take argtto be +πimmediately above the cut, and −πimmediately below it. This new definition is due to Hankel. Forzan integer, the cut is ineffective and we can close the contour to find 1 Γ(0)= 0;1 Γ(n)=1 (n−1)!, n> 0. (10.21) 10.1. THE GAMMA FUNCTION 405 Thus definitions A and B agree on the integers. It is less obvio us that they agree for all z. A hint that this is true stems integrating by parts 1 Γ(z)=1 2πi/bracketleftbigget (z−1)tz−1/bracketrightbigg−∞+i/epsilon1 −∞−i/epsilon1+1 (z−1)2πi/integraldisplay Cet tz−1dt=1 (z−1)Γ(z−1). (10.22) The integrated out part vanishes because etis zero at−∞. Thus the “new” gamma function obeys the same functional relation as the “ol d” one. To show the equivalence in general we will examine the definit ion B ex- pression for Γ(1−z) 1 Γ(1−z)=1 2πi/integraldisplay Cettz−1dt. (10.23) We will asume initially that Re z >0, so that there is no contribution from the small circle about the origin. We can therefore focus on c ontribution from the discontinuity across the cut 1 Γ(1−z)=1 2πi/integraldisplay Cettz−1dt=−1 2πi(2isinπ(z−1))/integraldisplay∞ 0tz−1e−tdt =1 πsinπz/integraldisplay∞ 0tz−1e−tdt. (10.24) The proof is then completed by using Γ( z)Γ(1−z) =πcosecπz,which we proved using definition A, to show that, under definition A, th e right hand side is indeed equal to 1 /Γ(1−z). We now use the uniqueness of analytic continuation, noting that if two analytic functions agree o n the region Re z> 0, then they agree everywhere. Infinite Product for Γ(z) The function Γ( z) has poles at z= 0,−1,−2,...therefore (zΓ(z))−1= (Γ(z+ 1))−1has zeros as z=−1,−2,.... Furthermore the integral in “defi- nition B” converges for all z, and so 1/Γ(z) has no singularities in the finite zplanei.e.it is an entire function. Thus means that we can use the infinit e product formula g(z) =g(0)ecz∞/productdisplay 1/braceleftbigg/parenleftbigg 1−z zj/parenrightbigg ez/zj/bracerightbigg (10.25) 406 CHAPTER 10. SPECIAL FUNCTIONS II for entire functions. We need to recall the definition of Euler-Mascheroni constan tγ=−Γ/prime(1) = .5772157..., and that Γ(1) = 1. Then 1 Γ(z)=zeγz∞/productdisplay 1/braceleftBig/parenleftBig 1 +z n/parenrightBig e−z/n/bracerightBig . (10.26) We can use this formula to compute 1 Γ(z)Γ(1−z)=1 (−z)Γ(z)Γ(−z)=z∞/productdisplay 1/braceleftBig/parenleftBig 1 +z n/parenrightBig e−z/n/parenleftBig 1−z n/parenrightBig ez/n/bracerightBig =z∞/productdisplay 1/parenleftbigg 1−z2 n2/parenrightbigg =1 πsinπz and so obtain another demonstration that Γ( z)Γ(1−z) =πcosecπz. Exercise 10.1 : Starting from the infinite product formula for Γ( z), show that d2 dz2ln Γ(z) =∞/summationdisplay n=01 (z+n)2. (Compare this “half series”, with the expansion π2cosec2πz=∞/summationdisplay n=−∞1 (z+n)2.) 10.2 Linear Differential Equations When a linear differential equation has meromorphic coeffeci ents, its solu- tions can be extended off the real line and into the complex pla ne. The broader horizon then allows us to see much more of their struc ture. 10.2.1 Monodromy Consider the linear differential equation Ly≡y/prime/prime+p(z)y/prime+q(z)y= 0, (10.27) 10.2. LINEAR DIFFERENTIAL EQUATIONS 407 wherepandqare meromorphic. Recall that the point z=ais aregular singular point of the equation if porqis singular there, but (z−a)p(z),(z−a)2q(z) (10.28) are both analytic at z=a. We know, from the explicit construction of power series solutions, that near a regular singular point yis a sum of functions of the formy= (z−a)αϕ(z) ory= (z−a)α(ln(z−a)ϕ(z) +χ(z)), where both ϕ(z) andχ(z) are analytic near z=a. We now examine this fact in a more topological way. Suppose that y1andy2are linearly independent solutions of Ly= 0. Start from some ordinary (non-singular) point of the equation and analytically continue the solutions round the singularity at z=aand back to the starting point. The continued functions ˜ y1and ˜y2will not in general coincide with the original solutions, but being still solutions of the equ ation, must be linear combinations of them. Therefore /parenleftbigg ˜y1 ˜y2/parenrightbigg =/parenleftbigg a11a12 a21a22/parenrightbigg/parenleftbigg y1 y2/parenrightbigg , (10.29) for some constants aij. By a suitable redefinition of the yiwe may either diagonalise this monodromy matrix to find /parenleftbigg ˜y1 ˜y2/parenrightbigg =/parenleftbigg λ10 0λ2/parenrightbigg/parenleftbigg y1 y2/parenrightbigg (10.30) or, if the eigenvalues coincide and the matrix is not diagona lizable, reduce it to a Jordan form /parenleftbigg ˜y1 ˜y2/parenrightbigg =/parenleftbigg λ1 0λ/parenrightbigg/parenleftbigg y1 y2/parenrightbigg . (10.31) These equations are satisfied, in the diagonalizable case, b y functions of the form y1= (z−a)α1ϕ1(z), y 2= (z−a)α2ϕ2(z), (10.32) whereλk=e2πiαk, andϕk(z) is single valued near z=a. In the Jordan-form case we must have y1= (z−a)α/bracketleftbigg ϕ1(z) +1 2πiλln(z−a)ϕ2(z)/bracketrightbigg , y 2= (z−a)αϕ2(z),(10.33) where again the ϕk(z) are single valued. Notice that coincidence of the monodromy eigenvalues λ1andλ2does not require the exponents α1andα2 408 CHAPTER 10. SPECIAL FUNCTIONS II to be the same, only that they differ by an integer. This is the s ame condition that signals the presence of a logarithm in the traditional s eries solution. The occurrence of fractional powers and logarithms in solut ions near a regular singular point is therefore quite natural. 10.2.2 Hypergeometric Functions Most of the special functions of Mathematical Physics are sp ecial cases of the hypergeometric function F(a,b;c;z), which may be defined by the series F(a,b;c;z) = 1 +a.b 1.cz+a(a+ 1)b(b+ 1) 2!c(c+ 1)z2+ +a(a+ 1)(a+ 2)b(b+ 1)(b+ 2) 3!c(c+ 1)(c+ 2)z3+···. =Γ(c) Γ(a)Γ(b)∞/summationdisplay 0Γ(a+n)Γ(b+n) Γ(c+n)Γ(1 +n)zn. (10.34) For general values of a,b,c, this series converges for |z|<1, the singularity restricting the convergence being a branch point at z= 1. Examples : (1 +z)n=F(−n,b;b;−z), (10.35) ln(1 +z) =zF(1,1; 2;−z), (10.36) z−1sin−1z=F/parenleftbigg1 2,1 2;3 2;z2/parenrightbigg , (10.37) ez= lim b→∞F(1,b; 1/b;z/b), (10.38) Pn(z) =F/parenleftbigg −n,n+ 1; 1;1−z 2/parenrightbigg , (10.39) where in the last line Pnis the Legendre polynomial. For future reference, note that expanding the right hand sid e as a powers series inzand integrating term by term shows that F(a,b;c;z) =Γ(c) Γ(b)Γ(c−b)/integraldisplay1 0(1−tz)−atb−1(1−t)c−b−1dt. (10.40) If Rec>Re (a+b), we may set z= 1 in this integral to get F(a,b;c; 1) =Γ(c)Γ(c−a−b) Γ(c−a)Γ(c−b). (10.41) 10.2. LINEAR DIFFERENTIAL EQUATIONS 409 The hypergeometric function is a solution of the second-ord er differential equation z(1−z)y/prime/prime+ [c−(a+b+ 1)z]y/prime−aby= 0. (10.42) this equation has regular singular points at z= 0,1,∞. Provided that 1 −c is not an integer, the general solution is y=AF(a,b;c;z) +Bz1−cF(b−c+ 1,a−c+ 1; 2−c;z). (10.43) The hypergeometric equation is a particular case of the gene ralFuchsian equation having three1regular singularities at z=z1,z2,z3. This equation is y/prime/prime+P(z)y/prime+Q(z)y= 0, (10.44) where P(z) =/parenleftbigg1−α−α/prime z−z1+1−β−β/prime z−z2+1−γ−γ/prime z−z3/parenrightbigg Q(z) =1 (z−z1)(z−z2)(z−z3)× /parenleftbigg(z1−z2)(z1−z3)αα/prime z−z1+(z2−z3)(z2−z1)ββ/prime z−z2+(z3−z1)(z3−z2)γγ/prime z−z3/parenrightbigg . (10.45) The parameters are subject to the constraint α+β+γ+α/prime+β/prime+γ/prime= 1, which ensures that z=∞is not a singular point of the equation. This 1The Fuchsian equation with tworegular singularities is y/prime/prime+p(z)y/prime+q(z)y= 0 with p(z) =/parenleftbigg1−α−α/prime z−z1+1 +α+α/prime z−z2/parenrightbigg q(z) =αα/prime(z1−z2)2 (z−z1)2(z−z2)2. Its general solution is y=A/parenleftbiggz−z1 z−z2/parenrightbiggα +B/parenleftbiggz−z1 z−z2/parenrightbiggα/prime . 410 CHAPTER 10. SPECIAL FUNCTIONS II equation is sometimes called Riemann’sP-equation . ThePprobably stands for Papperitz, who discovered it. The indicial equation relative to the regular singular poin t atz1is r(r−1) + (1−α−α/prime)r+αα/prime= 0, (10.46) and has roots r=α,α/prime. From this we deduce that Riemann’s equation has solutions which behave like ( z−z1)αand (z−z1)α/primenearz1. Similarly, there are solutions that behave like ( z−z2)βand (z−z2)β/primenearz2, and like (z−z3)γand (z−z3)γ/primenearz3. The solution space of Riemann’s equation is traditionally denoted by the Riemann “ P” symbol y=P  z1z2z3 α β γ z α/primeβ/primeγ/prime  (10.47) where the six quantities α,β,γ,α/prime,β/prime,γ/prime,are called the exponents of the so- lution. A particular solution is y=/parenleftbiggz−z1 z−z2/parenrightbiggα/parenleftbiggz−z3 z−z2/parenrightbiggγ F/parenleftbigg α+β+γ,α+β/prime+γ; 1 +α−α/prime;(z−z1)(z3−z2) (z−z2)(z3−z1)/parenrightbigg . (10.48) By permuting the triples ( z1,α,α/prime), (z2,β,β/prime), (z3,γ,γ/prime), and within them interchanging the pairs α↔α/prime,γ↔γ/prime, we may find a total2of 6×4 = 24 solutions of this form. They are called the Kummer solutions. Only two of these can be linearly independent, and a large part of the the ory of special functions is devoted to obtaining the linear relations betw een them. It is straightforward, but a trifle tedious, to show that (z−z1)r(z−z2)s(z−z3)tP  z1z2z3 α β γ z α/primeβ/primeγ/prime  =P  z1z2z3 α+r β +s γ +t z α/prime+r β/prime+s γ/prime+t   (10.49) providedr+s+t= 0. Riemann’s equation retains its form under M¨ obius maps, only the location of the singular points changing. We t herefore deduce that P  z1z2z3 α β γ z α/primeβ/primeγ/prime  =P  z/prime 1z/prime 2z/prime 3 α β γ z/prime α/primeβ/primeγ/prime  (10.50) 2The interchange β↔β/primeleaves the hypergeometric function invariant, and so does n ot give a new solution. 10.2. LINEAR DIFFERENTIAL EQUATIONS 411 where z/prime=az+b cz+d, z/prime 1=az1+b cz1+d, z/prime 2=az2+b cz2+d, z/prime 3=az3+b cz3+d. (10.51) By using the M¨ obius map which takes ( z1,z2,z3)→(0,1,∞), and by extracting powers to shift the exponents, we can reduce the g eneral eight- parameter Riemann equation to the three-parameter hyperge ometric equa- tion. ThePsymbol for the hypergeometric equation is F(a,b;c;z) =P  0∞ 1 0a 0z 1−c b c−a−b  . (10.52) Using this observation and a suitable M¨ obius map we see that F(a,b;a+b−c; 1−z) and (1−z)c−a−bF(c−b,c−a;c−a−b+ 1; 1−z) are also solutions of the Hypergeometric equation, each hav ing a pure (as opposed to a linear combination of) power-law behaviors nea rz= 1. (The previous solutions had pure power-law behaviours near z=0. ) These new solutions must be linear combinations of the old, and we may u se F(a,b;c; 1) =Γ(c)Γ(c−a−b) Γ(c−a)Γ(c−b),Re (c−a−b)>0, (10.53) together with the trick of substituting z= 0 andz= 1, to determine the coefficients and show that F(a,b;c;z) =Γ(c)Γ(c−a−b) Γ(c−a)Γ(c−b)F(a,b;a+b−c; 1−z) +Γ(c)Γ(a+b−c) Γ(a)Γ(b)(1−z)c−a−bF(c−b,c−a;c−a−b+ 1; 1−z). (10.54) This last equation holds for all values of a,b,csuch that the gamma functions make sense. 412 CHAPTER 10. SPECIAL FUNCTIONS II A complete set of pure-power solutions can be taken to be φ(0) 0(z) =F(a,b;c;z) φ(1) 0(z) =z1−cF(a+ 1−c,b+ 1−c; 2−c;z) φ(0) 1(z) =F(a,b; 1−c+a+b; 1−z) φ(1) 1(z) = (1−z)c−a−bF(c−a,c−b; 1 +c−a−b; 1−z) φ(0) ∞(z) =z−aF(a,a+ 1−c; 1 +a−b;z−1) φ(1) ∞(z) =z−bF(a,b+ 1−c; 1−a+b;z−1), (10.55) The connection coefficients are then φ(0) 0=Γ(c)Γ(c−a−b) Γ(c−a)Γ(c−b)φ(0) 1+Γ(c)Γ(a+b−c) Γ(a)Γ(b)φ(1) 1, φ(1) 0=Γ(2−c)Γ(c−a−b) Γ(1−a)Γ(1−b)φ(0) 1Γ(2−c)Γ(a+b−c) Γ(a+ 1−c)Γ(b+ 1−c)φ(1) 1, (10.56) and φ(0) 0=e−iπaΓ(c)Γ(b−a) Γ(c−a)Γ(b)φ(0) ∞+e−iπbΓ(2−c)Γ(a−b) Γ(a+ 1−c)Γ(1−b)φ(1) ∞, φ(1) 0=e−iπ(a+1−c)Γ(2−c)Γ(b−a) Γ(b+ 1−c)Γ(1−a)φ(0) ∞+e−iπ(b+1−c)Γ(2−c)Γ(a−b) Γ(a+ 1−c)Γ(1−b)φ(1) ∞. (10.57) These relations assume that Im z >0. The signs in the exponential factors must be reversed when Im z<0. Example: The P¨ oschel-Teller problem for general positive l.A substitution z= (1 +e2x)−1shows that the P¨ oschel-Teller Schrodinger equation /parenleftbigg −d2 dx2−l(l+ 1)sech2x/parenrightbigg ψ=Eψ (10.58) has solution ψ(x) = (1 +e2x)−κ/2(1 +e−2x)−κ/2F/parenleftbigg κ+l+ 1,κ−l;κ+ 1;1 1 +e2x/parenrightbigg , (10.59) 10.2. LINEAR DIFFERENTIAL EQUATIONS 413 whereE=−κ2. This solution behaves near x=∞as ψ∼e−κxF(κ+l+ 1,κ−l;κ+; 0) =e−κx. (10.60) We use the connection formula (10.54) to see that it behaves i n the vicinity ofx=−∞as ψ∼eκxF(κ+l+ 1,κ−l;κ+ 1; 1−e2x) →eκxΓ(κ+ 1)Γ(−κ) Γ(−l)Γ(1 +l)+e−κx Γ(κ+ 1)Γ(κ) Γ(κ+l+ 1)Γ(κ−l).(10.61) To find the bound-state spectrum, assume that κis positive. Then E=−κ2will be an eigenvalue provided that coefficient of e−κxnearx=−∞ vanishes. In other words, when Γ(κ+ 1)Γ(κ) Γ(κ+l+ 1)Γ(κ−l)= 0. (10.62) This condition is satisfied for a finite set κn,n= 1,...,[l] (where [l] denotes the integer part of l) at which κis positive but κ−lis zero or a negative integer. On setting κ=−ik, we find the scattering solution ψ(x) =/braceleftbigg eikx+r(k)e−ikxx/lessmuch0, t(k)eikxx/greatermuch0,(10.63) where r(k) =Γ(l+ 1−ik)Γ(−ik−l)Γ(ik) Γ(−l)Γ(1 +l)Γ(ik), =−sinπl πΓ(l+ 1−ik)Γ(−ik−l)Γ(ik) Γ(−ik), (10.64) and t(k) =Γ(l+ 1−ik)Γ(−ik−l) Γ(1−ik)Γ(−ik). (10.65) Wheneverlis a (positive) integer, the divergent factor of Γ( −l) in the de- nominator of r(k) causes the the reflected wave to vanish. This is something we had discovered in earlier chapters. In this particular ca se the transmission coefficientt(k) reduces to a phase t(k) =(−ik+ 1)(−ik+ 2)···(−ik+l) (−ik−1)(−ik−2)···(−ik−l). (10.66) 414 CHAPTER 10. SPECIAL FUNCTIONS II 10.3 Solving ODE’s via Contour integrals Our task in this section is to understand the origin of contou r integral solu- tions such as the expression F(a,b;c;z) =Γ(c) Γ(b)Γ(c−b)/integraldisplay1 0(1−tz)−atb−1(1−t)c−b−1dt, (10.67) we have previously seen for the hypergeometric equation. We are given a differential operator Lz=∂2 zz+p(z)∂z+q(z) (10.68) and seek a solution of Lzu= 0 as an integral u(z) =/integraldisplay ΓF(z,t)dt. (10.69) If we can find an Fsuch that LzF=∂Q ∂t, (10.70) for some function Q(z,t) then Lzu=/integraldisplay ΓLzF(z,t)dt=/integraldisplay Γ/parenleftbigg∂Q ∂t/parenrightbigg dt= [Q]Γ. (10.71) Thus, ifQvanishes at both ends of the contour, if it takes the same valu e at the two ends, or if the contour is closed and has no ends, we hav e succeeded in our quest. Example: Consider Legendre’s equation Lzu≡(1−z2)d2u dz2−2zdu dz+ν(ν+ 1)u= 0. (10.72) The identity Lz/braceleftbigg(t2−1)ν (t−z)ν+1/bracerightbigg = (ν+ 1)d dt/braceleftbigg(t2−1)ν+1 (t−z)ν+2/bracerightbigg (10.73) shows that Pν(z) =1 2πi/integraldisplay Γ/braceleftbigg(t2−1)ν 2ν(t−z)ν+1/bracerightbigg dt (10.74) 10.3. SOLVING ODE’S VIA CONTOUR INTEGRALS 415 will be a solution of Legendre’s equation provided that [Q]Γ≡/bracketleftbigg(t2−1)ν+1 (t−z)ν+2/bracketrightbigg Γ= 0. (10.75) We could, for example, take a contour that circles the points t=zandt= 1, but excludes the point t=−1. On going round this contour, the numerator aquires a phase of e2πi(ν+1), while the denominator of [ Q]Γaquires a phase of e2πi(ν+2). The net phase change is therefore e−2πi= 1. The function in the integrated-out part is therefore single-valued, and so the integrated-out part vanishes. When νis an integer, Cauchy’s formula shows that Pn(z) =1 2nn!dn dzn(z2−1)n, (10.76) which is Rodriguez’ formula for the Legendre polynomials. 1 −1 z Im Re(t) (t) Figure 10.2: Figure-of-eight contour for Qν(Z). The figure-of-eight contour shown in figure 10.2 gives us anot her solution Qν(z) =1 4isinπν/integraldisplay Γ/braceleftbigg(t2−1)ν 2ν(z−t)ν+1/bracerightbigg dt, ν /∈Z. (10.77) Here we define arg( t−1) and arg( t−1) to be zero for t>1. The integrated out part vanishes because the phase gained by the ( t2−1)ν+1in the numerator of [Q]Γduring the clockwise winding about t= 1 is undone during the anti- clockwise winding about t=−1, and, provided that zis outside the contour, there is no phase change in the ( z−t)−(ν+2)in the denominator. Whenνis real and positive the contributions from the circular arc s sur- roundingt=±1 become negligeable as we shrink this new contour down onto the real axis. After this manouvre the integral (10.77) becomes Qν(z) =1 2/integraldisplay1 −1/braceleftbigg(1−t2)ν 2ν(z−t)ν+1/bracerightbigg dt, ν > 0. (10.78) 416 CHAPTER 10. SPECIAL FUNCTIONS II In contrast to (10.77), this last formula continues to make s ense when ν is a positive integer, and so provides a convenient definitio n ofQn(z), the Legendre function of the second kind (See exercise 9.3). It is hard to find a suitable F(z,t) in one fell swoop. (The identity (10.73) exploited in the example is not exactly obvious!) An easier s trategy is to seek solution in the form of an integral operator with kernel Kacting on function v(t). Thus we set u(z) =/integraldisplayb aK(z,t)v(t)dt. (10.79) Suppose that LzK(z,t) =MtK(z,t), whereMtis differential operator in t that does not involve z. The operator Mtwill have have a formal adjoint M† t such that /integraldisplayb av(MtK)dt−/integraldisplayb aK(M† tv)dt= [Q(K,v)]b a. (10.80) (This is Lagrange’s identity.) Now Lzu=/integraldisplayb aLzK(z,t)vdt =/integraldisplayb a(MtK(z,t))vdt =/integraldisplayb aK(z,t)(M† tv)dt+ [Q(K,v)]b a. We can therefore solve the original equation, Lzu= 0, by finding a vsuch that (M† tv) = 0, and a contour with endpoints such that [ Q(K,v)]b a= 0. This may sound complicated, but an artful choice of Kcan make it much simpler than solving the original problem. Example : We will solve Lzu=d2u dz2−zdu dz+νu= 0, (10.81) by using the kernel K(z,t) =e−zt. We haveLzK(z,t) =MtK(z,t) where Mt=t2−t∂ ∂t+ν, (10.82) so M† t=t2+∂ ∂tt+ν=t2+ (ν+ 1) +t∂ ∂t. (10.83) 10.3. SOLVING ODE’S VIA CONTOUR INTEGRALS 417 The equation M† tv= 0 has solution v(t) =t−(ν+1)e−1 2t2, (10.84) and so u=/integraldisplay Γt−(1+ν)e−(zt+1 2t2)dt, (10.85) for some suitable Γ. 10.3.1 Bessel Functions As an illustration of the general method we will explore the t heory of Bessel functions. Bessel functions are member of the family of confluent hypergeo- metric functions , obtained by letting the two regular singular points z2,z3of the Riemann-Papperitz equation coalesce at infinity. The re sulting singular point is no longer regular, and confluent hypergeometric fun ctions have an essential singularity at infinity. The confluent hypergeome tric equation is zy/prime/prime+ (c−z)y/prime−ay= 0, (10.86) with solution Φ(a,c;z) =Γ(c) Γ(a)∞/summationdisplay n=0Γ(a+n) Γ(c+n)Γ(n+ 1)zn. (10.87) The second solution, when cis not an integer, is z1−cΦ(a−c+ 1,2−c;z). (10.88) We see that Φ(a,c;z) = lim b→∞F(a,b;c;z/b). (10.89) Other functions of this family are the parabolic cylinder functions , which in special cases reduce to e−z2/4times the Hermite polynomials , theerror function erf (z) =/integraldisplayz 0e−t2dt=zΦ/parenleftbigg1 2,3 2;−z2/parenrightbigg (10.90) and the Laguerre polynomials Lm n=Γ(n+m+ 1) Γ(n+ 1)Γ(m+ 1)Φ(−n,m+ 1;z). (10.91) 418 CHAPTER 10. SPECIAL FUNCTIONS II Bessel’s equation involves Lz=∂2 zz+1 z∂z+/parenleftbigg 1−ν2 z2/parenrightbigg . (10.92) Experience shows that a useful kernel is K(z,t) =/parenleftBigz 2/parenrightBigν exp/parenleftbigg t−z2 4t/parenrightbigg . (10.93) Then LzK(z,t) =/parenleftbigg ∂t−ν+ 1 t/parenrightbigg K(z,t) (10.94) soMis a first order operator, which is simpler to deal with than th e original second order Lz. In this case M†=/parenleftbigg −∂t−ν+ 1 t/parenrightbigg (10.95) and we need a vsuch that M†v=−/parenleftbigg ∂t+ν+ 1 t/parenrightbigg v= 0. (10.96) Clearlyv=t−ν−1will work. The integrated out part is [Q(K,v)]b a=/bracketleftbigg t−ν−1exp/parenleftbigg t−z2 4t/parenrightbigg/bracketrightbiggb a. (10.97) We see that Jν(z) =1 2πi/parenleftBigz 2/parenrightBigν/integraldisplay Ct−ν−1e“ t−z2 4t” dt. (10.98) solves Bessel’s equation provided we use a suitable contour . We can take for Ca contour starting at −∞−i/epsilon1and ending at−∞+i/epsilon1, and surrounding the branch cut of t−ν−1, which we take as the negative t axis. 10.3. SOLVING ODE’S VIA CONTOUR INTEGRALS 419 CRe(t) Im(t) Figure 10.3: Contour for solving Bessel equation. This contour works because Qis zero at both ends of the contour. A cosmetic rewrite t=uz/2 gives Jν(z) =1 2πi/integraldisplay Cu−ν−1ez 2(u−1 u)du. (10.99) Forνan integer, there is no discontinuity across the cut, so we ca n ignore it and takeCto be the unit circle. Then, recognizing the resulting Jn(z) =1 2πi/integraldisplay |z|=1u−n−1ez 2(u−1 u)du. (10.100) to be a Laurent coefficient, we obtain the familiar generating function ez 2(u−1 u)=∞/summationdisplay −∞Jn(z)un. (10.101) Whenνis not an integer, we see why we need a branch cut integral. If we setu=ewwe get Jν(z) =1 2πi/integraldisplay C/primedwezsinhw−νw, (10.102) whereC/primestarts goes from∞−iπto−iπ, to +iπto∞+iπ. 420 CHAPTER 10. SPECIAL FUNCTIONS II π π+i −iRe(w)Im(w) Figure 10.4: Bessel contour after change of variables. If we setw=t±iπon the horizontals and w=iθon the vertical part, we can rewrite this as Jν(z) =1 π/integraldisplayπ 0cos(νθ−zsinθ)dθ−sinνπ π/integraldisplay∞ 0e−νt−zsinhtdt. (10.103) All these are standard formulae for the Bessel function whos e origin would be hard to understand without the contour solutions trick. Whenνbecomes an integer, the functions Jν(z) andJ−ν(z) are no longer independent. In order to have a Bessel equation solution tha t retains its independence from Jν(z), even asνbecomes a whole number, we define the Neumann function Nν(z)def=Jν(z) cosνπ−J−ν(z) sinνπ =cotνπ π/integraldisplayπ 0cos(νθ−zsinθ)dθ−cosecνππ/integraldisplayπ 0cos(νθ+zsinθ)dθ −cosνπ π/integraldisplay∞ 0e−νt−zsinhtdt−1 π/integraldisplay∞ 0eνt−zsinhtdt. (10.104) 10.4. ASYMPTOTIC EXPANSIONS 421 +iπ π−iHνHν (2)(1) Figure 10.5: Contours defining H(1) ν(z)andH(2) ν(z). Both Bessel and Neumann functions are real for positive real x. Asxbecomes large they oscillate as slowly decaying sines and cosines. I t is sometimes convenient to decompose these real functions into solution s that behave as e±ix. We therefore define the Hankel functions by H(1) ν(z) =1 iπ/integraldisplay∞+iπ −∞ezsinhw−νwdw,|argz|<π/2 H(2) ν(z) =−1 iπ/integraldisplay∞−iπ −∞ezsinhw−νwdw,|argz|<π/2.(10.105) Then 1 2(H(1) ν(z) +H(2) ν(z)) =Jν(z), 1 2(H(1) ν(z)−H(2) ν(z)) =Nν(z). (10.106) 10.4 Asymptotic Expansions We often need the understand the behaviour of solutions of di fferential equa- tions and functions, such as Jν(x), whenxtakes values that are very large, or very small. This is the subject of asymptotics . 422 CHAPTER 10. SPECIAL FUNCTIONS II As an introduction to this art, consider the function Z(λ) =/integraldisplay∞ −∞e−x2−λx4dx. (10.107) Those of you who have taken a course quantum field theory based on path integrals will recognize that this is a “toy,” 0-dimensiona l, version of the path integral for the λϕ4model of a self-interacting scalar field. Suppose we wish to obtain the perturbation expansion for Z(λ) as a power series in λ. We naturally proceed as follows Z(λ) =/integraldisplay∞ −∞e−x2−λx4dx =/integraldisplay∞ −∞e−x2∞/summationdisplay n=0(−1)nλnx4n n!dx ?=∞/summationdisplay n=0(−1)nλn n!/integraldisplay∞ −∞e−x2x4ndx =∞/summationdisplay n=0(−1)nλn n!Γ(2n+ 1/2). (10.108) Something has clearly gone wrong here! The gamma function Γ( 2n+1/2)∼ (2n)!∼4n(n!)2overwhelms the n! in the denominator and the radius of convergence of the final power series is zero. The invalid, but popular, manoeuvre is the interchange of th e order of performing the integral and the sum. This interchange canno t be justified because the sum inside the integral does not converge unifor mly on the do- main of integration. Does this mean that the series is useles s? It had better not! All quantum field theory (and most quantum mechanics) pe rturbation theory relies on versions of this manoeuvre. We are saved to some (often adequate) degree because, while t he inter- change of integral and sum does not lead to a convergent serie s, it does lead to a valid asymptotic expansion . We write Z(λ)∼∞/summationdisplay n=0(−1)nλn n!Γ(2n+ 1/2) (10.109) where Z(λ)∼∞/summationdisplay n=0anλn(10.110) 10.4. ASYMPTOTIC EXPANSIONS 423 is shorthand for the more explicit Z(λ) =N/summationdisplay n=0anλn+O/parenleftbig λN+1/parenrightbig , N = 1,2,3,.... (10.111) The “bigO” notation Z(λ)−N/summationdisplay n=0anλn=O(λN+1) (10.112) asλ→0, means that lim λ→0/braceleftBigg |Z(λ)−/summationtextN 0anλn| |λN+1|/bracerightBigg =K <∞. (10.113) The basic idea is that, given a convergent power series/summationtext nanλnfor the functionf(λ), we fix the value of λand take more and more terms. The sum then gets closer to f(λ). Given an asymptotic expansion, on the other hand, we select a fixed number of terms in the series and then make λsmaller and smaller. The graph of f(λ) and the graph of our polynomial approximation then approach each other. The more terms we take the sooner th ey get close, but for any non-zero λwe can never get exacty f(λ)—no matter how many terms we take. We often consider asymptotic expansions where the independ ent variable becomes large. Here we have expansions in inverse powers of x: F(x) =N/summationdisplay n=0bnx−n+O/parenleftbig x−N−1/parenrightbig , N = 1,2,3.... (10.114) In this case F(x)−N/summationdisplay n=0bnx−n=O/parenleftbig x−N−1/parenrightbig (10.115) means that lim x→∞/braceleftBigg |F(x)−/summationtextN 0bnx−n| |x−N−1|/bracerightBigg =K <∞. (10.116) Again we take a fixed number of terms, and as xbecomes large the function and its approximation get closer. Observations: 424 CHAPTER 10. SPECIAL FUNCTIONS II i) Knowledge of the asymptotic expansion gives us useful kno wledge about the function, but does not give us everything. In particular , two distinct functions may have the same asymptotic expansion. For example, for small positive λ, the functions F(λ) andF(λ)+ae−b/λhave exactly the same asymptotic expansions as series in positive powers of λ. This is becausee−b/λgoes to zero faster than any power of λ, and so its asymp- totic expansion/summationtext nanλnhas every coefficient anbeing zero. Physicists commonly say that e−b/λis anon-perturbative function, meaning that it will not be visible to a perturbation expansion in powers o fλ. ii) An asymptotic expansion is usually valid only in a sector a<argz<b. Different sectors have different expansions. This is called t heStokes’ phenomenon . The most useful methods for obtaining asymptotic expansion s require that the function to be expanded be given in terms of an integr al. This is the reason why we have stressed the contour integral metho d of solving differential equations. If the integral can be approximated by a Gaussian, we are lead to the method of steepest descents . This technique is best explained by means of examples. 10.4.1 Stirling’s Approximation for n! We start from the integral representation of the Gamma funct ion Γ(z+ 1) =/integraldisplay∞ 0e−ttzdt (10.117) Sett=zζ, so Γ(z+ 1) =zz+1/integraldisplay∞ 0ezf(ζ)dζ, (10.118) where f(ζ) = lnζ−ζ. (10.119) We are going to be interested in evaluating this integral in t he limit that |z|→∞ and finding the first term in the asymptotic expansion of Γ( z+ 1) in powers of 1 /z. In this limit, the exponential will be dominated by the part of the integration region near the absolute maximum of f(ζ) Nowf(ζ) is a maximum at ζ= 1 and f(ζ) =−1−1 2(ζ−1)2+···. (10.120) 10.4. ASYMPTOTIC EXPANSIONS 425 So Γ(z+ 1) =zz+1e−z/integraldisplay∞ 0e−z 2(ζ−1)2+···dζ ≈zz+1e−z/integraldisplay∞ −∞e−z 2(ζ−1)2dζ =zz+1e−z/radicalbigg 2π z =√ 2πzz+1/2e−z. (10.121) By keeping more of the terms represented by the dots, and expa nding them as e−z 2(ζ−1)2+···=e−z 2(ζ−1)2/bracketleftbig 1 +a1(ζ−1) +a2(ζ−1)2+···/bracketrightbig ,(10.122) we would find, on doing the integral, that Γ(z+1)≈√ 2πzz+1/2e−z/bracketleftbigg 1 +1 12z+1 288z2−139 51840z3−571 24888320z4+O/parenleftbigg1 z5/parenrightbigg/bracketrightbigg . (10.123) Since Γ(n+ 1) =n! we also have n!≈√ 2πnn+1/2e−n/bracketleftbigg 1 +1 12n+···/bracketrightbigg . (10.124) We make contact with our discusion of asymptotic series by re writing the expansion as Γ(z+ 1)√ 2πzz+1/2e−z∼1 +1 12z+1 288z2−139 51840z3−571 24888320z4+...(10.125) This typical. We usually have to pull out a leading factor fro m the function whose asymptotic behaviour we are studying, before we are le ft with a plain asymptotic power series. 10.4.2 Airy Functions The Airy functions Ai( x) and Bi(x) are closely related to Bessel functions, and are named after the mathematician and astronomer George Biddell Airy. They occur widely in physics. We will investigate the behavi our of Ai(x) for 426 CHAPTER 10. SPECIAL FUNCTIONS II large values of|x|. A more sophisticated treatment is needed for this problem, and we will meet with Stokes’ phenomenon. Airy’s differentia l equation is d2y dz2−zy= 0. (10.126) On the real axis Airy’s equation becomes −d2y dx2+xy= 0, (10.127) and we we can think of this as the Schrodinger equation for a pa rticle running up a linear potential. A classical particle incident from th e left with total energyE= 0 will come to rest at x= 0, and then retrace its path. The point x= 0 is therefore called a classical turning point .The corresponding quantum wavefunction, Ai ( x), contains a travelling wave incident from the left and becoming evanescent as it tunnels into the classically forb idden region, x>0, together with a reflected wave returning to −∞. The sum of the incident and reflected waves is a real-valued standing wave. -10 -5 5 10 -0.4-0.20.20.4 Figure 10.6: The Airy function, Ai (x). We will look for contour integral solutions to Airy’s equati on of the form y(x) =/integraldisplay Cextf(t)dt. (10.128) Denoting the Airy differential operator by Lx≡∂2 x−x, we have Lxy=/integraldisplay C(t2−x)extf(t)dt=/integraldisplayb af(t)/braceleftbigg t2−d dt/bracerightbigg extdt. =/bracketleftbig −extf(t)/bracketrightbig C+/integraldisplay C/parenleftbigg/braceleftbigg t2+d dt/bracerightbigg f(t)/parenrightbigg extdt. (10.129) 10.4. ASYMPTOTIC EXPANSIONS 427 Thusf(t) =e−1 3t3and y(x) =/integraldisplayb aext−1 3t3dt. (10.130) The contour must end at points where the integrated-out term ,/bracketleftBig ext−1 3t3/bracketrightBig C, vanishes. There are therefore three possible contours, whi ch end at any two of +∞,∞e2πi/3,∞e−2πi/3. C1C C2 3 Figure 10.7: Contours providing solutions of Airy’s equation. Since the integrand is an entire function, the sum yC1+yC2+yC3is zero, so only two of the three solutions are linearly independent. Th e Airy function itself is defined by Ai (z) =1 2πi/integraldisplay C1ext−1 3t3dt=1 π/integraldisplay∞ 0cos/parenleftbigg xs+1 3s3/parenrightbigg ds (10.131) In obtaining last equality, we have deformed the contour of i ntegration, C1, that ran from∞e−2πi/3to∞e2πi/3so that it lies on the imaginary axis, and there we have written t=is. You may check ( ` a laJordan) that this deformation does not alter the value of the integral. To study the asymptotics of this function we need to examine s eparately two casesx/greatermuch0 andx/lessmuch0. For both ranges of x, the principal contribution to the integral will come from the neighbourhood of the stati onary points off(t) =xt−t3/3. These stationary points are never pure maxima or 428 CHAPTER 10. SPECIAL FUNCTIONS II minima of the real part of f(the real part alone determines the magnitude of the integrand) but are always saddle points . We must deform the contour so that on the integration path the stationary point is the hi ghest point in a mountain pass. We must also ensure that everywhere on the contour the difference between fand its maximum value stays real. Because of the orthogonality of the real and imaginary part contours, this means that we must take a path of steepest descent from the pass — hence the name of the method. If we stray from the steepest descent path, the ph ase of the exponent will be changing. This means that the integrand wil l oscillate and we can no longer be sure that the result is dominated by the con tributions near the saddle point. b) a) uv v u Figure 10.8: Steepest descent contours and location and orientation of t he saddle passes for a) x/greatermuch0, b)x/lessmuch0. i)x/greatermuch0 : The stationary points are at t=±√x. Writingt=ξ−√xhave f(ξ) =−2 3x3/2+ξ2√x−1 3ξ3(10.132) while neart= +√xwe writet=ζ+√xand find f(ζ) =−2 3x3/2−ζ2√x−1 3ζ3(10.133) We see that the saddle point near −√xis a local maximum when we route the contour vertically, while the saddle point near +√xis a local maximum as we go down the real axis. Since the contour in Ai( x) is 10.4. ASYMPTOTIC EXPANSIONS 429 aimed vertically we can distort it to pass through the saddle point near −√x, but cannot find a route through the point at +√xwithout the integrand oscillating wildly. At the saddle point the expon ent,xt−t3/3, is real. If we write t=u+ivwe have Im (xt−t3/3) =v(x−u2+v3/3), (10.134) so the exact steepest descent path, on which the imaginary pa rt remains zero is given by the union of real axis ( v= 0) and the curve u2−1 3v2=x. (10.135) This is a hyperbola, and the branch passing through the saddl e point at−√xis plotted in a). Now setting ξ=is, we find Ai (x) =1 2πe−2 3x3/2/integraldisplay∞ −∞e−√xs2+···ds∼1 2√πx−1/4e−2 3x3/2.(10.136) ii)x/lessmuch0 : The stationary points are now at ±i/radicalbig |x|. Settingt=ξ±i/radicalbig |x| find that f(x) =∓i2 3|x|3/2∓iξ2/radicalbig |x|. (10.137) The exponent is no longer real, but the imaginary part will be constant and the integrand non-oscillatory provided we deform the co ntour so that it becomes the disconnected pair of curves shown in b). T he new contour passes through both saddle points and we must sum their contributions. Near t=i/radicalbig |x|we setξ=e3πi/4sand get 1 2πie3πi/4e−i2 3|x|3/2/integraldisplay∞ −∞e−√ |x|s2ds=1 2i√πe3πi/4|x|−1/4e−i2 3|x|3/2 =−1 2i√πe−iπ/4|x|−1/4e−i2 3|x|3/2 (10.138) Neart=−i/radicalbig |x|we setξ=e2πi/3sand get 1 2iπeπi/4ei2 3|x|3/2/integraldisplay∞ −∞e−√ |x|s2ds=1 2i√πeπi/4|x|−1/4ei2 3|x|3/2(10.139) 430 CHAPTER 10. SPECIAL FUNCTIONS II The sum of these two contributions is Ai (x)∼1√π|x|1/4sin/parenleftbigg2 3|x|3/2+π 4/parenrightbigg . (10.140) The fruit of our labours is therefore Ai (x)∼1 2√πx−1/4e−2 3x3/2/bracketleftbigg 1 +O/parenleftbigg1 x/parenrightbigg/bracketrightbigg , x> 0, ∼1√π|x|1/4sin/parenleftbigg2 3|x|3/2+π 4/parenrightbigg/bracketleftbigg 1 +O/parenleftbigg1 x/parenrightbigg/bracketrightbigg , x< 0. (10.141) Suppose that we allow xto become complex x→z=|z|eiθ, with−π < θ < π . Then figure 10.9 shows how the steepest contour evolves and l eads the two quite different expansion for positive and negative x. We see that for 0< θ < 2π/3 the steepest descent path continues to be routed through the single stationary point at −/radicalbig |z|eiθ/2. Onceθreaches 2π/3, though, it passes through both stationary points. The contribution to the integral from the newly aquired stationary point is, however, expone ntially smaller as|z|→∞ than that of t=−/radicalbig |z|eiθ/2. The new term is therefore said to besubdominant , and makes an insignificant contribution to the asymptotic behaviour of Ai ( z). The two saddle points only make contributions of the same magnitude when θreachesπ. If we analytically continue beyond θ=π, the new saddlepoint will now dominate over the old, and only i ts contribtion is significant at large |z|. The Stokes line , at which we must change the form of the asymptotic expansion is therefore at θ=π. If we try to systematically keep higher order terms we will fin d, for the oscillating Ai (−z), a double series Ai (−z)∼π−1/2z−1/4/bracketleftBigg sin(ρ+π/4)∞/summationdisplay n=0(−1)nc2nρ−2n −cos(ρ+π/4)∞/summationdisplay n=0(−1)nc2n+1ρ−2n−1/bracketrightBigg (10.142) whereρ= 2z3/2/3. In this case, therefore we need to extract two leading coefficients before we have asymptotic power series. The subject of asymptotics contains many subtleties, and th e reader in search of a more detailed discussion is recommened to read Be nder and Orszags Advanced Mathematical methods for Scientists and Engineer s. 10.4. ASYMPTOTIC EXPANSIONS 431 -2 -1 0 1 2-2-1012 -2 -1 0 1 2-2-1012 -2 -1 0 1 2-2-1012 -2 -1 0 1 2-2-1012a) b) c) d) Figure 10.9: Evolution of the steepest-descent contour from passing thr ough only one saddle point to passing through both. The dashed and solid lines are contours of the real and imaginary parts, repectively, of (zt−t3/3).θ= Argz takes the values a) 7π/12, b)15π/24, c)2π/3, d)9π/12. 432 CHAPTER 10. SPECIAL FUNCTIONS II Exercise 10.2 : Consider the behaviour of Bessel functions when xis large. By applying the method of steepest descent to the Hankel functi on contours show that H(1) ν(x)∼/radicalbigg 2 πxei(x−νπ/2−π/4)/bracketleftbigg 1−4ν2−1 8πx+···/bracketrightbigg H(2) ν(x)∼/radicalbigg 2 πxe−i(x−νπ/2−π/4)/bracketleftbigg 1 +4ν2−1 8πx+···/bracketrightbigg , and hence Jν(x)∼/radicalbigg 2 πx/bracketleftbigg cos/parenleftBig x−νπ 2−π 4/parenrightBig −4ν2−1 8xsin/parenleftBig x−νπ 2−π 4/parenrightBig +···/bracketrightbigg , Nν(x)∼/radicalbigg 2 πx/bracketleftbigg sin/parenleftBig x−νπ 2−π 4/parenrightBig +4ν2−1 8xcos/parenleftBig x−νπ 2−π 4/parenrightBig +···/bracketrightbigg . 10.5 Elliptic Functions The subject of elliptic functions goes back to remarkable id entities of Guilio Fagnano (1750) and Leonhard Euler (1761). Euler’s formula i s /integraldisplayu 0dx√ 1−x4+/integraldisplayv 0dy/radicalbig 1−y4=/integraldisplayr 0dz√ 1−z4, (10.143) where 0≤u,v≤1, and r=u√ 1−v4+v√ 1−u4 1 +u2v2. (10.144) This looks mysterious, but perhaps so does /integraldisplayu 0dx√ 1−x2+/integraldisplayv 0dy/radicalbig 1−y2=/integraldisplayr 0dz√ 1−z2, (10.145) where r=u√ 1−v2+v√ 1−u2, (10.146) until you realize that the latter formula is merely sin(a+b) = sinacosb+ cosasinb (10.147) 10.5. ELLIPTIC FUNCTIONS 433 in disguise. To see this set u= sina, v = sinb (10.148) and remember the integral formula for the inverse trig funct ion a= sin−1u=/integraldisplayu 0dx√ 1−x2. (10.149) The Fagnano-Euler formula is a similarly disguised additio n formula for an elliptic function . Just as we use the substitution x= sinyin the 1/√ 1−x2 integral, we can use an elliptic function substitution to ev aluate elliptic in- tegrals such as I4=/integraldisplayx 0dt/radicalbig (t−a1)(t−a2)(t−a3)(t−a4)(10.150) I3=/integraldisplayx 0dt/radicalbig (t−a1)(t−a2)(t−a3). (10.151) The integral I3is a special case of I4, wherea4has been sent to infinity by use of a M¨ obius map t→t/prime=at+b ct+d, dt/prime= (ad−bc)dt (ct+d)2. (10.152) Indeed, we can use a suitable M¨ obius map to send any three of t he four pointsanto 0,1,∞. The idea of elliptic functions (as opposed to the integrals, which are their functional inverse) was known to Gauss, but Abel and Jacobi w ere the first to publish (1827). For the general theory, the simplest elli ptic function is the Weierstrass ℘. This is defined by first selecting two linearly independent periodsω1,ω2, and setting ℘(z) =1 z2+/summationdisplay (m,n)/negationslash=0/braceleftbigg1 (z−mω1−nω2)2−1 (mω1+nω2)2/bracerightbigg .(10.153) The sum is over integers m,n, positive and negative, but not both 0. Helped by the counterterm, the sum is absolutely convergent, so we c an rearrange the terms to prove double periodicity ℘(z+mω1+nω2) =℘(z). (10.154) 434 CHAPTER 10. SPECIAL FUNCTIONS II The function is thus determined everywhere by its values in t he period paral- lelogramP={λω1+µω2: 0≤λ,µ< 1}. Double periodicity is the defining characteristic of elliptic functions. .. .... .. ωω2 xy 1.. Figure 10.10: Unit cell and double-periodicity. Any non-constant meromorphic function, f(z), which is doubly periodic has four basic properties: a) The function must have at least one pole in its unit cell. Ot herwise it would be holomorphic and bounded, and therefore a constan t by Liouville. b) The sum of the residues at the poles must add to zero. This fo llows from integrating f(z) around the boundary of the period parallelogram and observing that the contributions from opposite edges ca ncel. c) The number of poles in each unit cell must equal the number o f zeros. This follows from integrating f/prime/fround the boundary of the period parallelogram. d) Iffhas zeros at the Npointsziand poles at the Npointspithen N/summationdisplay i=1zi−N/summationdisplay i=1pi=nω1+mω2 wherem,nare integers. This follows from integrating zf/prime/fround the boundary of the period parallelogram. The Weierstass ℘has a second-order pole at the origin. It also obeys lim |z|→0/parenleftbigg ℘(z)−1 z2/parenrightbigg = 0, 10.5. ELLIPTIC FUNCTIONS 435 ℘(z) =℘(−z), ℘/prime(z) =−℘/prime(−z). (10.155) The property that makes ℘(z) useful for evaluating integrals is (℘/prime(z))2= 4℘3(z)−g2℘(z)−g3, (10.156) where g2= 60/summationdisplay (m,n)/negationslash=01 (mω1+nω2)4, g 3= 140/summationdisplay (m,n)/negationslash=01 (mω1+nω2)6.(10.157) Equation (10.156) is proved by examining the first few terms i n the Laurent expansion in zof the difference of the left hand and right hand sides. All negative powers cancel, as does the constant term. The differ ence is zero at z= 0, has no poles or other singularities, and being continuou s and periodic is automatically bounded. It is therefore identically zero by Liouville’s theorem. From the symmetry and periodicity of ℘we see that ℘/prime(z) = 0 atω1/2, ω2/2 and (ω1+ω2)/2 where℘(z) takes values e1=℘(ω1/2),e2=℘(ω2/2), ande3=P((ω1+ω2)/2). Now℘/primemust have exactly three zeros since it has a pole of order three at the origin and, by property c), the numb er of zeros in the unit cell is equal to the number of poles. We therefore kno w the location of all three zeros and can factorize 4℘3(z)−g2℘(z)−g3= 4(℘−e1)(℘−e2)(℘−e3). (10.158) We note that the coefficient of ℘2in the polynomial on the left side is zero, implying that e1+e2+e3= 0. This is consistent with property d). The rootseican never coincide. For example, ( ℘(z)−e1) has a double zero atω1/2, but two zeros is all it is allowed because the number of pole s per unit cell equals the number of zeros, and ( ℘(z)−e1) has a double pole at 0 as its only singularity. Thus ( ℘−e1) cannot be zero at another point, but it would be if e1coincided with e2ore3. As a consequence, the discriminant ∆ = 16(e1−e2)2(e2−e3)2(e1−e3)2=g3 2−27g2 3, (10.159) is never zero. We use℘to write z=℘−1(u) =/integraldisplayu ∞dt 2/radicalbig (t−e1)(t−e2)(t−e3)=/integraldisplayu ∞dt/radicalbig 4t3−g2t−g3. (10.160) 436 CHAPTER 10. SPECIAL FUNCTIONS II This maps the uplane cut from e1toe2ande3to∞one-to-one onto the 2-torus, regarded the unit cell of the ωn,m=nω1+mω2lattice. Aszsweeps over the torus, the points x=℘(z),y=℘/prime(z) move on the elliptic curve y2= 4x3−g2x−g3 (10.161) which should be thought of as a set in CP2. These curves, and the finite fields of rational points that lie on them, are exploited in modern c ryptography. The magic which leads to addition formula, such as the Euler- Fagnano relation with which we began this section, lies in the (not im mediatley obvi- ous) fact that any elliptic function having the same periods as℘(z) can be expressed as a rational function of ℘(z) and℘/prime(z). From this it follows (after some thought) that any two such elliptic functions, f1(z) andf2(z), obey a relationF(f1,f2) = 0, where F(x,y) =/summationdisplay an,mxnym(10.162) is a polynomial in xandy. We can eliminate ℘/prime(z) in these relations at the expense of introducing square roots. modular invariance Ifω1andω2are periods and define a unit cell, so are ω/prime 1=aω1+bω2 ω/prime 2=cω1+dω2 wherea,b,c,d are integers with ad−bc=±1. This condition on the deter- minant ensures that the matrix inverse also has integer entr ies, and so the ωi can be expressed in terms of the ω/prime iwith integer coefficients. Consequently the set of integer linear combinations of the ω/prime igenerate the same lattice as the integer linear combinations of the original ωi. This notion of redefining the unit cell should be familiar to your from solid state phys ics. If we wish to preserve the orientation of the basis vectors, we must res trict ourselves to maps whose determinant ad−bcis unity. The set of such transforms constitute the the modular group SL(2 ,Z). Clearly℘is invariant under this group, as are g2andg3and ∆. Now define ω2/ω1=τ, and write g2(ω1,ω2) =1 ω4 1,˜g2(τ), g 3(ω1,ω2) =1 ω6 1,˜g3(τ).∆(ω1,ω2) =1 ω12 1˜∆(τ), (10.163) 10.5. ELLIPTIC FUNCTIONS 437 and also J(τ) =˜g3 2 ˜g3 2−27˜g2 3=˜g3 2 ˜∆. (10.164) Because the denominator is never zero when Im τ >0, the function J(τ) is holomorphic in the upper half-plane — but not on the real axis . The function J(τ) is called the elliptic modular function . Except for the prefactors ωn 1, the functions ˜ gi(τ),˜∆(τ) andJ(τ) are invariant under the M¨ obius transformation τ→aτ+b cτ+d. (10.165) with /parenleftbigg a b c d/parenrightbigg ∈SL(2,Z). (10.166) This M¨ obius transformation does not change if the entries i n the matrix are multiplied by a common factor of ±1, and so the transformation is an element of the modular group PSL(2 ,Z)≡SL(2,Z)/{I,−I}. Taking into account the change in the ωα 1prefactors we have ˜g2/parenleftbiggaτ+b cτ+d/parenrightbigg = (cτ+d)4˜g3(τ), ˜g3/parenleftbiggaτ+b cτ+d/parenrightbigg = (cτ+d)6˜g3(τ), ˜∆/parenleftbiggaτ+b cτ+d/parenrightbigg = (cτ+d)12˜∆(τ). (10.167) Becausec= 0 andd= 1 for the special case τ→τ+1, these three functions obeyf(τ+1)−f(τ) and so depend on τonly via the combination q2=e2πiτ. For example, it is not hard to prove that ˜∆(τ) = (2π)12q2∞/productdisplay n=1/parenleftbig 1−q2n/parenrightbig24. (10.168) We can also expand them as power series in q2— and here things get interest- ing because the coefficients have number-theoretic properti es. For example ˜g2(τ) = (2π)4/bracketleftBigg 1 12+ 20∞/summationdisplay n=1σ3(n)q2n/bracketrightBigg , ˜g3(τ) = (2π)6/bracketleftBigg 1 216−7 3∞/summationdisplay n=1σ5(n)q2n/bracketrightBigg . (10.169) 438 CHAPTER 10. SPECIAL FUNCTIONS II The symbol σk(n) is defined by σk(n) =/summationtextdkwheredruns over all positive divisors of the number n. In the case of the function J(τ), the prefactors cancel and J/parenleftbiggaτ+b cτ+d/parenrightbigg =J(τ), (10.170) soJ(τ) is amodular invariant . One can show that if J(τ1) =J(τ2),then τ2=aτ1+b cτ1+d(10.171) for some modular transformation with integer a,b,c,d , wheread−bc= 1, and further, that any modular invariant function is a ration al function of J(τ). It seems clear that J(τ) is rather a special object. ThisJ(τ) is the function referred to on page 174 in connection with th e Monster group. As with the ˜ gi,J(τ) depends on τonly through q2. The first few terms in the power series expansion of J(τ) in terms of q2turn out to be 1728J(τ) =q−2+744+196884 q2+21493760q4+864299970 q6+···.(10.172) SinceAJ(τ)+Bhas all the same modular invariance properties as J(τ), the numbers 1728 = 123and 744 are just conventional normalizations. Once we set the coefficient of q−2to unity, however, the remaining integer coefficients are completely determined by the modular properties. A numb er-theory interpretation of these integers seemed lacking until John McKay and others observed that that 1 = 1 196884 = 1 + 196883 21493760 = 1 + 196883 + 21296786 864299970 = 2 ×1 + 2×196883 + 21296786 + 842609326 , (10.173) where “1” and the large integers on the right-hand side are th e dimensions of the smallest irreducible representations of the Monster. T his “Monstrous Moonshine” was originally mysterious and almost unbelieva ble, (“moon- shine” = “fantastic nonsense”) but it was explained by Richa rd Borcherds by the use of techniques borrowed from string theory.3Borcherds received the 1998 Fields Medal for this work. 3“I was in Kashmir. I had been traveling around northern India , and there was one 10.6. FURTHER EXERCISES AND PROBLEMS 439 10.6 Further Exercises and Problems Exercise 10.3 : Show that the binomial series expansion of (1 + x)−νcan be written as (1 +x)−ν=∞/summationdisplay m=0(−x)mΓ(m+ν) Γ(ν)m!,|x|<1. Exercise 10.4 :A Mellin transform and its inverse . Combine the Beta-function identity (10.15) with a suitable change of variables to eval uate the Mellin transform /integraldisplay∞ 0xs−1(1 +x)−νdx, ν > 0, of (1 +x)−νas a product of Gamma functions. Now consider the integral 1 2πiΓ(ν)/integraldisplayc+i∞ c−i∞x−sΓ(ν−s)Γ(s)ds. Here Rec∈(0,ν). The contour therefore runs parallel to the imaginary axis with the poles of Γ( s) to its left and the poles of Γ( ν−s) to its right. Use the identity Γ(s)Γ(1−s) =πcosecπs to show that when |x|<1 the contour can be closed by a large semicircle lying to the left of the imaginary axis. By using the preceding exer cise to sum the contributions from the enclosed poles at s=−n, evaluate the integral. Exercise 10.5 :Mellin-Barnes integral . Use the technique developed in the preceding exercise to show that F(a,b,c;−x) =Γ(c) 2πiΓ(a)Γ(b)/integraldisplayc+i∞ c−i∞x−sΓ(a−s)Γ(b−s)Γ(s) Γ(c−s)ds, for a suitable range of x. This integral representation of the hypergeometric function is due to the English mathematician Ernest Barnes ( 1908), later a controversial Bishop of Birmingham. really long tiresome bus journey, which lasted about 24 hour s. Then the bus had to stop because there was a landslide and we couldn’t go any further. It was all pretty darn unpleasant. Anyway, I was just toying with some calculation s on this bus journey and finally I found an idea which made everything work”- Richard B orcherds (Interview in The Guardian , August 1998). 440 CHAPTER 10. SPECIAL FUNCTIONS II Exercise 10.6 : Let Y=/parenleftbiggy1 y2/parenrightbigg Show that the matrix differential equation d dxY=A zY+B 1−zY, where A=/parenleftbigg0a 0 1−c/parenrightbigg , B =/parenleftbigg0 0 b a+b−c+ 1/parenrightbigg , has a solution Y(z) =F(a,b,;c,z)/parenleftbigg1 0/parenrightbigg +z aF/prime(a,b;c;z)/parenleftbigg0 1/parenrightbigg . Exercise 10.7 :Kniznik-Zamolodchikov equation. The monodromy properties of solutions of differential equations play an important rol e in conformal field theory. The Fuchsian equations studied in this exercise are obeyed by the correlation functions in the level- kWess-Zumino-Witten model. LetV(a),a= 1,...n, be spin-jarepresentation spaces for the group SU(2). Let W(z1,...,zn) be a function taking values in V(1)⊗V(2)⊗···⊗V(n). (In other wordsWis a function Wi1,...,in(z1,...,zn) where the index ialabels states in the spin-jafactor.) Suppose that Wobeys the Kniznik-Zamolodchikov (K-Z) equations (k+ 2)∂ ∂zaW=/summationdisplay b,b/negationslash=aJ(a)·J(b) za−zbW, a = 1,...,n, where J(a)·J(b)≡J(a) 1J(b) 1+J(a) 2J(b) 2+J(a) 3J(b) 3, andJ(a) iindicates the su(2) generator Jiacting on the V(a)factor in the tensor product. If we set z1=z, for example and fix the position of z2,...zn, then the differential equation in zhas regular singular points at the n−1 remaining zb. a) By diagonalizing the operator J(a)·J(b)show that there are solutions W(z) that behave for zaclose tozbas W(z)∼(za−zb)∆j−∆ja−∆jb, where ∆j=j(j+ 1) k+ 2,∆ja=ja(ja+ 1) k+ 2, 10.6. FURTHER EXERCISES AND PROBLEMS 441 andjis one of the spins |ja−jb|≤j≤j1+jaoccuring in the decompo- sition ofja⊗jb. b) Define covariant derivatives ∇a=∂ ∂za−/summationdisplay b,b/negationslash=aJ(a)·J(b) za−zb and show that [∇a,∇b] = 0. Conclude that the effect of parallel transport of the solutions of the K-Z equations provides a representat ion of the braid group of the world lines of the za. 442 CHAPTER 10. SPECIAL FUNCTIONS II Index p-chain, 125, 305 p-cycle, 305 p-form, 48 addition theorem for elliptic functions, 433 Airy’s equation, 426 algebraic geometry, 12 analytic signal, 379 anti-derivation, 50 atlas, 34 Bargmann, Valentine, 310 Bergman space, 310 Bergman, Stefan, 310 Bernoulli numbers, 384 Berry’s phase, 263 Beta function, 403 Betti number, 116, 128, 345 Bianchi identity, 69 Bochner Laplacian, 169 Bogomolnyi equation, 110 Borcherds, Richard, 438 Borel-Weil-Bott theorem, 265 boundary conditions Dirichlet, Neumann and Cauchy, 298 branch cut, 341 branch point, 341 branching rules, 200, 251Brouwer degree, 89, 160 bulk modulus, 22 bundle co-tangent, 58 tangent, 34 trivial, 258 vector, 34 Calugareanu relation, 103 Cartan algebra, 246 Cartan, ´Elie, 37, 224 Casimir operator, 239 Cayley’s theorem for groups, 176 chain complex, 127 chart, 34 Christoffel symbols, 64 Cicero, Marcus Tullius, 85 closed form, 51, 59 co-ordinates Cartesian, 18 conformal, seeco-ordinates, isother- mal isothermal, 349 co-root vector, 247 cohomology, 121 commutator, 40 complex algebraic curve, 345 complex differentiable, 293 complex projective space, 12, 91 443 444 INDEX constraint holonomic versus anholonomic, 43 contour, 303 Cornu spiral, 368 covector, 2 cup product, 142 curl as a differential form, 51 d’Angelo, John, 296 D-bar problem, 308 Darboux co-ordinates, 60, 61, 267 theorem, 59 de Rham’s theorem, 140 de Rham, Georges, 121 degree-genus relation, 346 derivation, 45, 53 derivative complex, 293 convective, 112 covariant, 63 exterior, 49, 50 Lie, 45 descent equations, 288 diffeomorphism, 116 dimensional regularization, 331 Dirac gamma matrices, 227 dispersion relation, 372 distribution involutive, 42 of tangent fields, 41 distributions principal part, 367 domain, 294 elliptic function, 344, 433elliptic modular function, 437 embedding, 347 entire function, 326, 333 equivalence relation, 174 essential singularity, 326, 333 Euler angles, 43, 70, 220 character, 131, 158, 345 class, 153 Euler-Maclaurin sum formula, 384 Euler-Mascheroni constant, 406 exact form, 51 exact sequence, 131 long, 136 short, 133, 136 exponential map, 216 Fermat’s liittle theorem, 176 Feynman path integral, 100 fibre, 257 fibre bundle, 39 field covector, 37 tangent vector, 35 flow incompressible, 294 irrotational, 294 of tangent vector field, 40 foliation, 41 form closed, 59 Fredholm operator, 157 Fresnel integrals, 368 Frobenius’ integrability theorem, 42 reciprocity theorem, 206 Frobenius-Schur indicator, 204 INDEX 445 Gauss linking number, 100 Gauss-Bonnet theorem, 153, 284 Gauss-Bruhat decomposition, 392 Gell-Mann “ λ” matrices, 242 generating function for Chern character, 151 genus, 345 geometric phase, seeBerry’s phase geometric quantization, 265 gradient as a covector, 37 Grassmann, Herman, 14 Green, George, 25 Haar measure, 230 harmonic conjugate, 294 Hilbert transform, 378 Hodge “⋆” map, 55, 350 decomposition, 157 theory, 154 Hodge, William, 154 homeomorphism, 116 homology group, 127 homotopy, 96, 227 class, 96 Hopf bundle, seemonopole bundle index, 98, 223 map, 94, 220, 222 horocycles, 354 ideal, 235 immersion, 347 index theorem, 158, 390, 393 induced metric, 83 induced representation, 205infinitesimal homotopy relation, 53 interior multiplication, 53 intersection form, 144 Jacobi identity, 60, 234 Jordan form, 407 Killing field, 46 form, 236 Killing, William, 46 Kramer’s degeneracy, 211 Lagrange’s theorem, 174 Lam´ e constants, 22 Laplace-Beltrami operator, 156 Laplacian acting on vector field, 154 Legendre function, 374 Legendre function Qn(x), 415 Levi-Civita symbol, 17 Lie algebra, 207 bracket, 40, 234 derivative, 45 Lie, Sophus, 207 line bundle, 258 Lipshitz’ formula, 384 Lobachevski geometry, 110, 354 M¨ obius strip, 258 manifold, 34 orientable, 79 Riemann, 66 map anti-conformal, 298 isogonal, 298 modular group, 436 446 INDEX monodromy, 406 monople bundle, 279 monopole bundle, 265 moonshine, monstrous, 174, 438 Morse function, 159 Morse index theorem, 160 multilinear form, 11 M¨ obius map, 339, 433 Neumann’s formula, 374 Nyquist criterion, 375 orbit,of group action, 178 order of group, 172 orientable manifold, 78 P¨ oschel-Teller equation, 412 pairing, 2, 138 Pauliσmatrices, 93, 211 period and de Rham’s theorem, 140 of elliptic function, 344 Peter-Weyl theorem, 231 Pfaffian system, 44 Pl¨ ucker relations, 16, 31 Pl¨ ucker, Julius, 16 Plemelj formulæ, 372 Poincar´ e disc, 110, 354 duality, 159 lemma, 50, 117 Poincar´ e-Hopf theorem, 160 Poisson bracket, 60 Poisson’s ratio, 23 pole, 308 Pontryagin class, 153 principal bundle, 257principal part integral, 364 product cup, 142 direct, 181 group axioms, 171 tensor, 10 wedge, 13, 49 projective plane, 129 quaternions, 211 quotient group, 174 space, 179 rank of Lie algebra, 246 of tensor, 5 residue, 308 resolution of the identity, 190 retraction, 117 Riemann Psymbol, 410 sum, 304 surface, 341 Rodriguez’ formula, 415 rolling conditions, 43, 107 root vector, 244 Russian formula, 288 section, 259 of bundle, 39 Serret-Frenet relations, 107 sextant, 224 shear modulus, 22 sheet, 341 simplex, 122 simplicial complex, 123 Skyrmion, 91 space INDEX 447 homogeneous, 179 retractable, 117 spinor, 93, 224 stereographic map, 92 Stokes’ line, 430 phenomenon, 424 theorem, 84 strain tensor, 48 stream-function, 295 streamline, 295 structure constants, 214 symplectic form, 59 tangent bundle, 34 space, 33 tantrix, 103 tensor Cartesian, 18 curvature, 66 isotropic, 19 strain, 20, 48 stress, 20 torsion, 66 theorem Blasius, 317 Darboux, 59 de Rham, 140 Frobenius integrability, 42 Frobenius’ reciprocity, 206 Gauss-Bonnet, 153, 284 Lagrange, 174 Morse index, 160 Peter-Weyl, 231 Picard, 333 Poincar´ e-Hopf, 160 residue, 308Riemann mapping, 300 Stokes, 84 Theta function, 337 topological current, 98 torsion in homology, 130 of curve, 107 tensor, 66 transfom Hilbert, 378 variety, 12 Segre, 12 vector bundle, 63 Laplacian, 154 velocity potential, 294 vielbein, 64 orthonormal, 69, 148 volume form, 84 Weierstrass ℘function, 433 weight, 243 Weitzenb¨ ock formula, 168 Weyl’s identity, 210 Wiener-Hopf sum equations, 387 winding number, 89 Young’s modulus, 23