Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / blackbody boltzmann

blackbody

DOCX · 87.4 KB
Open DOCX file

Exploratory notes by Phil (dated 12.20.14, with red annotations added 6.2.15) on what the classical blackbody model means. They review thermal equilibrium, phase space and Liouville's theorem, and Reif's and Zemansky's treatments of the partition function. The main part counts accessible states, applies Stirling's approximation, and sets up maximizing Ω under fixed particle number and energy, with Phil's later note that the unconstrained derivative step was wrong.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
The Blackbody Problem PhL 12.20.14 [ this doc sat idle for 6 months, I went off and learned about the Lagrange multiplier method, then I reread this on 6.2.15 and added red notes starting on that date.] What exactly was the classical model of the "black body radiator" -- a furnace that perhaps glows red hot inside, and you sample the radiation coming out with some instruments. The first idea that comes to mind is that there is some kind of "thermal equilibrium" between the furnace walls and the EM fields in the cavity, but it is not clear what this really means. I suspect there are many pieces to this blackbody scenario which will take some time to put together. Notion of a State We really do need some concept of thermal equilibrium in order to say anything at all about the blackbody problem. So let's start with a review of Baby Reif statistical mechanics notes made in 2007. The general framework here is a "system of particles" such as a dilute gas in a box. Right off the bat, Reif is talking about such a system being in a "state" and is talking about the "density of states near energy E". Classically the "state" of such a system (assume particles all have same mass m and do not interact etc etc) might be specified by giving the position qi and velocity of vi of each particle. But the "density of states near energy E" is a bit unclear. My Reif notes at this point seem to assume quantum mechanics and I am off computing the density of states for a particle in a 1D box. For the blackbody problem, I might be off doing the QM calculation of states for a harmonic oscillator. But let's try to stay classical as long as possible. After all, it is the blackbody failure that motivates QM, so I don't want to assume QM a priori. I have dim memories of "phase space" for a system of classical particles and products dpidqi and the possible hint that the density of states in phase space is constant. I want to continue through the Reif notes, but first I want to have some handle on the meaning of the density of states for a classical system. I think you can talk about statistical mechanics in a world with no quantum mechanics. Web sites mention something called Liouville's Theorem which seems to say that the density of states in phase space is somehow constant, whatever that means. Goldstein has a short discussion of this theorem on page 266 as an application of Poisson brackets. Here is a comment from the wiki phase space page which I quote since it mentions some concepts of interest Statistical ensembles in phase space The motion of an ensemble of systems in this space is studied by classical statistical mechanics. The local density of points in such systems obeys Liouville's Theorem, and so can be taken as constant. Within the context of a model system in classical mechanics, the phase space coordinates of the system at any given time are composed of all of the system's dynamic variables. Because of this, it is possible to calculate the state of the system at any given time in the future or the past, through integration of Hamilton's or Lagrange's equations of motion. So OK, I will now try to learn more about this subject which I have no memory of ever dealing with before. Zemansky's thermo book has a stat mech section on page 251, but like Reif, right off the bat he is talking about particle in a box and quantum numbers and quantum states. The key idea is that there exists some discrete spectrum of states with energies he calls εi . Z's presentation seems very good and very general without making use of any particular quantum problem. He gets right to the idea that Ni = gie-βε where Ni is the number of particles in state i which has degeneracy gi and energy εi. Z goes on as Reif does to show that the unknown constant β is related to "temperature" and Z = Σi gie-βε is the "partition function". But a classical era physicist would never be using a discrete εi spectrum based on quantum theory, so how did such a person approach this subject? Boltzmann is credited on the Z function Ludwig Eduard Boltzmann (February 20, 1844 – September 5, 1906) was an Austrian physicist and philosopher whose greatest achievement was in the development of statistical mechanics, which explains and predicts how the properties of atoms (such as mass, charge, and structure) determine the physical properties of matter (such as viscosity, thermal conductivity, and diffusion). I read quite a bit in the Stanford history, but it never quite states what I want. But the Drexel piece and others do "start off" with an interesting starting point, namely http://www.pages.drexel.edu/~cfa22/msim/node6.html Here the exponent is e-βH where H is the system energy expressed in terms of the phase space coordinates that Hamiltonians always use. And there is the phase space integral, and there is no weighting function involve, the density in phase space is just "1", as commented above (as a conjecture). For me right now, the important idea is that there are no quantized energies εi floating around, we just have H as a continuous function of the x and p variables. It seems that the symbol Q is often used for the partition function, although Reif and Zemansky both use Z. One could just regard the above general form as an ansatz. Boltzmann (Stanford) liked the simple form e-βH because for a single particle it factors into a term for each dimension, and others point out that it is invariant under a certain coordinates transformation which matches the unity Jacobian of the phase space under such a transformation. But we really are faced with a question here: why is the probability of a point in phase space reduced as H increases? It might be good to look at the above mentioned derivation. This is in a book written in 2002 which I see is not very accessible. The Lagrange Multiplier Approach of Zemansky and The Origin of the Expo Factor Perhaps one can accept the discrete state approach and then take a continuum limit to end up with a derivation of the "Boltzmann factor" e-βE . Then we can describe a "system of particles" where for each particle there is a spectrum of energy states εi and state i will hold Ni particles. We can allow (as in Zemansky page 255) that an energy state εi has degeneracy gi which just means there are gi available states at energy εi . These are common QM concepts, but let's just take them to be a step on the way to the continuum analysis. Suppose we have N1 particles existing in state 1. For each of these N1 particles, there were g1 equivalent states, so there are g1N1 ways to put the N1 particles in state 1 which has energy ε1. Why is this? Well, take the first ball and there are g1 ways to put it into one of N1 boxes. Then take the second ball and there are g1 ways to put it into one of N1 boxes. That explains the g1N1 . For example, one way might be a00bc where the ball labeled a goes into the first box, etc. Then b00ac would be another way. But you can see that if the balls are all labeled a, then these ways are a00aa and a00aa and you have thus overcounted. In fact, you have overcounted by a factor N1, BUT, this correction is only correct if you ignore configurations with more than one ball per box, so this is where the assumption that g1 >> N1 sneaks in! So with this assumption if the particles are identical, then the correct number of different ways to put the N1 particles into the g1 degenerate states of energy ε1 is going to be g1N1/ N1! Now if you have N1 particles at energy ε1, and N2 at energy ε2 and so on. Then you get this factor for each of your energy levels. This then leads to Zemansky (10-4) and p 253-255 which says Ω = .... = number of "accessible states" for the full set of particles Here we are assuming a finite number m of available states with energies ε1, ε2.....εm . Notice then that we have made two assumptions: (1) the energy levels are discrete; (2) there are a finite number of energy levels. For the QM hydrogen atom, (1) is true but (2) is not true, for example. Now think of Ω(N1, N2.....Nm) = Ω(N) as a function of the Ni and suppose this function has a very strong peak at some specific vector value N = J, say. Then if we have an ensemble of such particle systems, and if we represent each member of the ensemble by its value of N (its ensemble state), then we would expect those values of N to cluster close to the peak value J. The systems are not all exactly at J, but they are all close to J. The reason is merely that J is statistically the most likely value for each system's N due to the assumed "strong peak". So where would such a strong peak come from? To find out, lets take the ln of the above expression: lnΩ = ( N1lng1 - ln N1! ) + ( N2lng2 - ln N2! ) + .... + ( Nmlngm - ln Nm! ) = Σi=1m [ Ni ln gi ] - Σi=1m [ ln Ni! ] The first term makes Ω large, the second term makes it small. If we now make the assumption that for a true system of interest all the Ni are "large numbers", we can use Stirling that ln x! = x ln x - x. Zemansky gives a graphical explanation on page 256, but I am happy with the result. Then we have lnΩ = Σi=1m [ Ni ln gi ] - Σi=1m [ Ni ln Ni - Ni ] = Σi=1m [ Ni ln ] + Σi=1m Ni = Σi=1m [ Ni ln ] + M // Zemansky (10-6) where M is the total number of particles. So now we are assuming also that all of the systems in our ensemble have exactly M particles. I don't yet see why there should be a sharp peak in ln Ω or Ω. But now we go on with another assumption: each of the ensembles has the same total energy U. Then we have two "conditions" which can be stated in this simple manner Σi=1m Ni = M Σi=1m Ni εi = U where M and U are constants. So here then is the problem set up for the Lagrange Multiplier Method. The idea is to maximize the function Ω subject to the two constraints: Ω(N1,N2...Nm) = .... = number of "accessible states" for the full set of particles Σi=1m Ni = M // number of particles is a constant (so not photons I guess) Σi=1m Ni εi = U // total energy is a constant (adiabatic I guess) When I first wrote this doc, I did not know how to solve this kind of problem, now I think I do. Now if we are looking for a peak in Ω or lnΩ, we should set all the derivatives to zero. [ When there are constraints, this is the wrong thing to do! Thus, much below is based on a wrong premise ] Then 0 = ∂ lnΩ/∂Ni for i = 1,2....m This would indicate a simultaneous peak in all the Ni variables, not sure such a thing exists. [ not with the constraints] Now ∂x [ x ln ] = ∂x [ x ln g - x ln x ] = ln g - ∂x [x ln x] = ln g - { x * 1/x + ln x } = ln g - {1 + ln x } = ln (g/x) - 1 So then ∂Nj [ Ni ln ] = [ ln (gi/Ni) - 1 ] δi,j and then 0 = ∂Nj lnΩ = ∂Nj[ Σi=1m [ Ni ln ] + M ] j = 1,2...m so m equations here = ∂Nj[ Σi=1m { Ni ln } ] = Σi=1m ∂Nj{ Ni ln } = Σi=1m [ δi,j ln (gi/Ni) - 1 ] = ln (gj/Nj) - 1 Thus I have shown : ∂Nj lnΩ = ln (gj/Nj) - 1 Note: If it is true that gj >> Nj, then we have a large positive slope coming to our peak coming in from any direction j. This makes it seem to be a strong peak. Now we can write d [ln Ω ] = Σj=1m ∂Nj lnΩ * dNj = Σj=1m [ ln (gj/Nj) - 1] * dNj = Σj=1m ln (gj/Nj)dNj - Σi=jm dNj = Σj=1m ln (gj/Nj)dNj - dM = Σi=1m ln (gi/Ni) dNi (10-9) where in the last line we have used the condition Σi=1m Ni = M = constant so dM = 0. Also in the last line we replace dummy index j by i. I still don't see any strong peak emerging yet. So far, here is what we know d [ln Ω ] = 0 since we are at a simultaneous peak in all variables Think of just two variables N1 and N2 as being (x,y) and we plot ln Ω of (x,y). At a peak, we expect both derivatives to vanish, just think of the plot, a smooth hilltop. Really we also have N3 but it is determined from the sum M, so only N1 and N2 are independent variables. So we now have three equations Σi=1m ln (gi/Ni) dNi = 0 // this says d [ln Ω ] = 0 Σi=1m dNi= 0 // this says dM = 0 Σi=1m εidNi = 0 // this says dU = 0 Now comes the magic step. Multiply line 2 by lnA and line 4 by -β where A and β are unknown constants, and I guess we assume neither is 0 : Σi=1m ln (gi/Ni) dNi = 0 // this says d [ln Ω ] = 0 Σi=1m ln A dNi= 0 // this says dM = 0 - Σi=1m β εidNi = 0 // this says dU = 0 Now add the three equations to get Σi=1m [ ln (gi/Ni) + ln A - β εi] dNi = 0 Think of this equation as a vector equation which says F dN = 0 where dN is a variation in N in some direction in N space done at the assumed peak. Since dN can be in any direction, in order to have this equation be true, you really need F = 0. I am not being totally rigorous here, but it seems reasonable. We conclude then that ln (gi/Ni) + ln A - β εi = 0 or ln (A gi/Ni) = β εi or A gi/Ni= eβε or Ni / (Agi) = e-βε or Ni = A gi e-βε (10-10) This is the fascinating result!! The "energy exponential" has just "appeared" in our result. Where "did it come from" ? It came from solving the following mathematical problem: " Given m accessible states i = 1,2...m having degeneracy gi and energy εi, and given fixed M total identical particles (where we assume M is very large), and assuming the particles having fixed total energy U, how would you insert the particles into the states in order to cause the function Ω to be a maximum? " A couple of loose ends here. How do we know it is a max and not a min? ∂Nj lnΩ = ln (gj/Nj) - 1 // first derivative, really m different equations ∂Nj2 lnΩ = ∂Nj [ ∂Nj lnΩ] = ∂Nj [ln (gj/Nj) - 1] = ∂Nj [ln (gj/Nj)] = ∂Nj [ ln gj - ln Nj] = - ∂Nj ln Nj = - 1/Nj < 0 Thus in all directions the curvature is "cupping down", so yes, we really have a maximum. If we sum (10-10) over i we get M = A Σi gi e-βε ≡ A Z Z = Σi gi e-βε where Z is the famous partition function. Then A = M/Z and we can write Ni = M ( gi e-βε) / Z (10-14) I still don't see why the peak is a strong peak! The curvature in the jth direction is - 1/Nj and if Nj is large, then the curvature is in fact small, so the peak is NOT strong as I thought it was. Well consider this: ∂Nj lnΩ = (1/Ω) ∂NjΩ so that ∂NjΩ = Ω ∂Nj lnΩ and then ∂Nj2Ω = ∂Nj [ Ω ∂Nj lnΩ ] = ∂NjΩ * ∂Nj lnΩ + Ω * ∂Nj2 lnΩ = [ Ω ∂Nj lnΩ ] * ∂Nj lnΩ + Ω * ∂Nj2 lnΩ = Ω [∂Nj lnΩ ]2 + Ω * ∂Nj2 lnΩ = Ω [ ln (gj/Nj) - 1 ]2 + Ω * [ - 1/Nj] = Ω { [ ln (gj/Nj) - 1 ]2 - 1/Nj } ≈ Ω { [ ln (gj/Nj) - 1 ]2 } where we neglect the -1/Nj relative to the +1 . Now our solution was ln (gi/Ni) + ln A - β εi = 0 ln (gi/Ni) = β εi - lnA so then ∂Nj2Ω ≈ Ω { [ β εi - lnA - 1 ]2 } Well, perhaps the argument is that no matter what the size is of the second factor, Ω is a monstrous number and that is why the curvature is huge and why the peak is so strong. Too bad that Zemansky did not comment on this. On page 253 Zemansky claims that for actual systems, one always has gi >> Ni and he shows this for the 3D particle in box system. I guess I don't see why this is generally true. I do see that gi is sort of the volume of an octant shell of some radius and this shell has a lot of Planck boxes in it. Maybe the particle occupies only one of these boxes. In any event, if gi >> Ni then ln (gj/Nj) >> 1 as well and then above I have shown that ∂Nj2Ω ≈ Ω [ ln (gj/Nj) ]2 Now how would you convert the above discussion to a continuum of energy levels? Go back first to, Ω = .... = number of "accessible states" for the full set of particles This is not easy to deal with. For one thing, it could be some kind of infinite product since m → 0, but I really don't know how to write that. However, consider instead, lnΩ = Σi=1m ln → ∫ dε ln where now we are integrating over a continuous energy spectrum ε. At any ε, there is some g(ε) and some N(ε) density of particles. But this is rather painful still. Suppose we stick with discrete as in the above until we get to this result Ni = A gi e-βε Z = Σi gi e-βε The sum is over "states". For a classical particle, we could denote a state by a point in phase space. We could quantize phase space artificially by breaking it in to little areas of size h0n where n is the dimensionality of the Cartesian space. Then we interpret Σi as a sum over these blocks, and let that become an integral to get Z = (1/h0n) ∫∫ dnp dnq g(p,q) e-βE(p,q) n = 3 for 3D space You could write E(p,q) = H(p,q) = the Hamiltonian for the particle. Again, this is all for a single particle and p,q is a point in phase space for such a particle. BUT this formulation does not allow for interactions between particles. It is as if the εi had nothing to do with relative positions of particles, but if we are trying to apply this to a gas of interacting particles, then you cannot write E = H(p,q) for one particle, because probably H depends on the qi of all the particles. So it is really not clear how you apply the formalism above to a gas of particles. If the particles do not interact, then perhaps the above makes sense. But in that case I suppose we have only H(p) and then Z = (1/h0n) V ∫ dnp e-βH(p) where V is the volume of space in n dimensions. Main conclusion of this section. I see how it is that we get these important results Ni = A gi e-βε = M gi e-βε / Z where Z = Σi gi e-βε = partition function In the derivation we assumed that the states were discrete and so could be enumerated by a sum. We assumed that the energies εi were fixed and not affected by particle interactions. We assumed a fixed number of particles M and degeneracies gi. We did not assume the validity of "quantum mechanics" except that the states could be enumerated (countable number of states). We assumed that the occupation numbers Ni were large integers to justify the Stirling formula, and we assumed identical particles. The key new result for me is how the "Boltzmann exponential factor" comes out naturally from the posed math problem and is not something that must be assumed. It is not clear to me exactly how you would take the continuum limit for describing the states, but it can probably be done somehow. I do realize that for a gas of interacting particles, if you identify the state energies εi with the Hamiltonian, you end up with a contradiction since then in effect the εi are not really constants, but the above derivation assumed they were in fact constants. Example. Suppose there are m = 3 states with these energies ε1 = 1 ε2 = 2 ε3 = 3 The two constraint equations are these, let us say, N1 + N2 + N3 = 100 1 = (1,1,1)/ = n1/ ε1N1 + ε2N2 + ε3N3 = 40 or N1 + 2N2 + 3N3 = 40 2 = (1,2,3)/= n2/ Each constraint is represented by a plane with the normal shown and passing distances 10 and 20 from the origin at closest approach. The two constraint equations are really (this all from my Ng notes) n1 N = 100 => 1 N = 100 => 1 N = 100/ n2 N = 40 => 2 N = 40 => 2 N = 40/ Now the intersection of two planes is a line in 3-space, so that line is the locus of possible solutions. There will I think be a single point on that line segment (in quadrant 1) for which Ω takes on a maximal value. I don't offhand know the equation of the line. Now, one way to "solve" this problem is to let N3 be a "variable" and solve for N1 and N2 like so N1 + N2 + N3 = 100 // first as is N1 + 2N2 + 3N3 = 40 // second as is Now subtract first from second to get N2 + 2N3 = - 60 => N2 = - 60 - 2N3 Stop. This says N2 is always negative! Then put this into the first to get N1 + (-60 - 2N3) + N3 = 10 => N1 - N3 = 70 => N1 = N3 + 70 Now here is the Ω formula Ω = where we can assume whatever degeneracies we like. We then have Ω =Ω(N3) = Obviously N3 can at most run from 0 to 10, so we can just plot the above function. Suppose for example we had g1 = 10 g2 = 20 g3 = 20. Then we just let Maple do the plot. But I have to get reasonable numbers, so start over again like this: N1 + N2 + N3 = M e1N1 + e2N2 + e3N3 = U Have maple just solve this for N1 and N2: Write this out N2 = [ -e1*N3 +e1*M +e3*N3 - U] / (e1-e2) N1 = [ -e3*N3 -e2*M +e2*N3 + U] / (e1-e2) Assume that e1 > e2. In order to have both N1 and N2 be positive we need to have -e1*N3 +e1*M +e3*N3 - U > 0 -e3*N3 -e2*M +e2*N3 + U > 0 or (e3-e1)*N3 +e1*M - U > 0 (e2-e3)*N3 -e2*M + U > 0 Add these to find that (e2-e1)N3 - (e2-e1)M > 0 or (e1-e2)N3 - (e1-e2)M < 0 or N3 - M < 0 or N3 < M which is a bit of a no brainer. Maybe rewrite the pair this way (e3-e1)*N3 > U - e1*M (e2-e3)*N3 > -U + e2*M Now suppose we assume U is small and M is large. Then roughly we have (e3-e1)*N3 > - e1*M (e2-e3)*N3 > + e2*M or (e3/e1-1)*N3 > - M (1-e3/e2)*N3 > + M or (e3/e1-1) > - M/N3 (1-e3/e2) > + M/N3 or (1 - e3/e1) < M/N3 (1-e3/e2) > M/N3 or (1 - e3/e1) < M/N3 < (1-e3/e2) The right < seems to require that (1-e3/e2) be positive, which means 1 > e3/e2 . Suppose we assume e2 > 0. Then we end up requiring that e2 > e3 and e1 > e2, which then says e3 < e2 < e1 e2 > 0 This seems totally strange to me! I am obviously missing some main idea, as per usual. Day over. ****************** items of interest **************** http://www4.ncsu.edu/~franzen/public_html/CH795N/lecture/XIII/XIII.html http://plato.stanford.edu/entries/statphys-Boltzmann/ http://www.pages.drexel.edu/~cfa22/msim/node6.html