Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Calculus, Real Analysis, Topology

more about e

DOCX · 449.9 KB
Open DOCX file

Expository note by Phil dated 9.2.09, written after reading Ahlfors, covering real variables only. It argues that a^x has no meaning for non-integer x until defined, requires a^x a^y = a^(x+y), derives f'(x) = f'(0)f(x), and develops ln and exp properties from the integral definition. It concludes a^x = exp(x ln a) and derives power rules; the later part was not seen.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
More about e PhL 9.2.09 I have been reading Ahlfors and then looked back at my own "about e" document, and I am still a bit confused about the "batting order" for knowing facts in this area. My comments here relate only to real variables, not complex variables. See Ahlfors page 43 for the extension to complex argument. Question : What is the meaning of ax and why is exp(x) = ex ? The second part of the question is this: How do we know that we can write the function exp(x), which we know is the inverse of the ln(x) function, as "some constant " "raised to the power x" ? As will be seen below, the whole concept of raising any constant a to a non-integer power x, as in ax, has no meaning at all until we give it a meaning! At the start, ax only "has meaning" for us if x = a positive integer n. What is the value, for example, of 3π ? How would you compute this value? We show below that it is possible to "continue" the meaning of an to non-integer exponents x in such a way that ax takes the correct values at x = n, and also provides a smooth interpolation between these integer values. (1) When x = n = 1,2,3... we know the meaning: an = a*a*..*a n times. (2) We also know that for integer values n and m, anam = an+m and also (an)m = anm. For example, 2223 = 2*2 2*2*2 = 2*2*2*2*2 = 25. And (22)3 = (2*2) (2*2) (2*2) = 2*2*2*2*2*2 = 26. (3) We know nothing about a1/2 at this point. Same for a1/n. Same for arbitrary ax. For example, at this point we don't know that (ax)y = axy except when x and y are integers, so we cannot say (a1/2)2 = a. Again, our problem is that the entity ax is not even defined yet for non-integer x. Yes, we could make a special definition for a1/n as the nth root of an, and thus extend the meaning of ax to x = integers and inverse integers, but let's NOT do this, it will come out of our interpolation anyway. (4) We attempt to define a new function, never seen before, which is written this way: f(x) = ax where x is a real number and a is some real constant. We want this new function f(x) to somehow be a "continuation" off the integer values x = n. That is, we want to be sure that f(2) = a2 = a*a and similarly that f(5) = a5 = a*a*a*a*a. So what we seek is a function ax which "interpolates" between the an values. (5) There are an infinite number of such interpolating functions. For example, if we had one such function f1(x), we could create many more by adding Ansin(nπx) for arbitrary An coefficient sets. This added piece of course vanishes at all the positive integers. (6) We shall now attempt to nail down a particular interpolating function by requiring that the function have this property: axay = ax+y for any real pair x and y. In other words, f(x)f(y) = f(x+y). This requirement is not unreasonable, since we already know it is true for x and y both integers, as in (2) above. Our requirement concerning our particular interpolating f(x) of interest then has the following corollary implications: [ note that 6.1 says f(x) is a representation of the "addition group" ] axay = ax+y "requirement" 6.1 f(x)f(y) = f(x+y) ax1ax2....axn = ax1+x2+...+xn (6a) f(x1)f(x2)...f(xn) = f(x1+x2+...xn) (ax)n = anx (6b) [f(x)]n = f(nx) axa-x = ax-x = a0 (6c) f(x)f(-x) = f(0) At this point, we don't know what a0 means. If we set x = 0 in (6c) we get a0a0 = a0. This suggests that we make the additional requirement concerning our interpolating function that a0 = f(0) = 1. We are of course allowed to add all the requirements we want, but we have to make sure that any solution interpolating function we finally arrive at meets all these requirements. Suppose in item (6b) we select x = 1/n. We then get (a1/n)n = a1 = a (6d) [f(1/n)]n = f(1) Now for the first time we have an interpretation of f(1/n) = a1/n . It is the number which you multiply by itself n times to get a. Thus, we now know that a1/n = "the nth root of a". We have in mind that a is some positive real number, so there is no problem with this thing being true. So this fact has fallen out of our approach, although we don't yet know how to fully define ax yet. That is, we know that when we finally come up (below) with a formula for computing (defining) ax (which formula meets our requirements above), then that formula will correctly give us the roots of numbers. Status to this point: We are looking for an interpolating function f(x) = ax which satisfies axay = ax+y and which satisfies a0 = 1. For x = n, we know that f(n) = an = a*a...*a n times, these are the values we are trying to interpolate. We also have a meaning for a1/n for n a positive integer. (7) Let's now look at the derivative of f(x) = ax, using informal notation: [f(x+dx) - f(x)]/ dx = [ax+dx - ax]/ dx = ax [ adx - 1]/dx // used 6.1 in the last step We have assumed that f(0) = a0 = 1, so let's attempt a linear fit Taylor series so f(dx) = adx ≈ f(0) + dx f '(0) + ... = 1 + dx f '(0) so adx - 1 ≈ dx f '(0) . Thus, for small dx we have [ adx - 1]/dx = f '(0). Our derivative line above then reads, f '(x) = ax f '(0) = f(x) f '(0) We have thus arrived at a new fact concerning f(x): it's derivative is itself times a constant ! Moreover, we can write the above as df(x)/ f(x) = f '(0) dx => !Syntax Error, Idf/f = f '(0) x 7.l (8) At this point we digress to review and in fact derive many properties of the ln(x) and exp(x) functions. From the first definition below, the reader sees a connection to 7.1 which we will resume later in (9). Our general goal is to make some connection between f(x) = ax and these ln(x) and exp(x) functions. We actually derive more facts than we "need" but we want to get them down once and for all. We start with this definition of the so-called ln(x) function: (Thomas p 390) ln(x) ≡ !Syntax Error, Idt/t just the area under this curve 1/x starting at 1 from which we conclude at once that ∂x ln(x) = 1/x 8.1 ln(1) = 0 8.2 From our knowledge of first order ODE's, we can draw this conclusion: Fact 1: the function ln(x) is the unique solution of this ODE system: F '(x) = 1/x with F(1) = 0 Since ln(x) is a monotonic function, we can easily define its inverse which we call the exp(x) function: exp(x) ≡ ln-1(x) exp(ln(x)) = x ln(x) = exp-1(x) ln(exp(x)) = x Setting x = 1 in the 2nd equation we get 1 = exp(ln(1)) = exp(0). Differentiating the 4th equation we get (1/exp(x)) ∂x(exp(x)) = 1. We can rewrite these last two equations as ∂x(exp(x)) = exp(x) 8.3 exp(0) = 1 8.4 which says the exp(x) function is its own derivative (!) and takes the value 1 at zero argument. From our knowledge of first order ODE's, we can draw this conclusion: Fact 2: the function exp(x) is the unique solution of this ODE system: G '(x) = G(x) with G(0) = 1. Now consider this application of the chain rule, then of 8.1 with x→xy, then finally of 8.1 as is: ∂x ln(yx) = [∂xy ln(yx) ] y = [ 1/(xy)] y = 1/x = ∂x ln(x) This says that the functions ln(yx) and ln(x) have the same derivative for all x, so we conclude that ln(yx) = ln(x) + Cy Setting x = 1, we find that ln(y) = ln(1)+Cy = Cy, and thus we have proven that ln(xy) = ln(x) + ln(y) 8.5 and repeating this n times we easily show that (later in 10.1 we extend this to arbitrary real n) ln(xn) = n ln(x) 8.6 These facts are certainly not obvious looking only at the definition ln(x) ≡ !Syntax Error, Idt/t . Meanwhile, using 8.5 again we have 0 = ln(1) = ln(x/x) = ln(x) + ln(1/x) from which we conclude ln(1/x) = - ln(x) 8.7 Then setting y = 1/z in 8.5 we get ln(x/z) = ln(x) - ln(z) so we have then shown that ln(x/y) = ln(x) – ln(y) 8.8 Next, turning to the exp(x) function, consider this fact: ∂x[ exp(x) exp(a-x)] = exp(x) exp(a-x) (-1) + exp(x) exp(a-x) = 0 chain rule Therefore, exp(x) exp(a-x) = Ca. Set x = 0 so that exp(0) exp(a) = Ca so Ca = exp(a). Set y = a-x and we conclude that exp(x) exp(y) = exp(x+y) 8.9 // Ahlfors p 43 credit which of course tells us that exp(x) exp(-x) = exp(0) = 1 so that exp(-x) = 1/exp(x) 8.10 Notice at once that of we define h(x) = exp(x), 8.9 says h(x)h(y) = h(x+y), which reminds us of 6.1 above where we imposed the requirement on our function f(x) that f(x)f(y) = f(x+y). So we think we may be on the right track here, approaching a solution to the problem we set out to solve. (9) We now apply our section (8) knowledge to our situation in item (7) and we arrive at this conclusion, starting with our equation 7.1 above: [ here just using definitions of ln and exp ] !Syntax Error, Idf/f = f '(0) x => ln[f(x)] = f '(0) x => f(x) = exp(f '(0) x ) If we set x = 1 in the middle equation, we find f '(0) = ln[f(1)] = ln(a1) = ln(a) So we have a our sought-after "interpolating function" as follows: f(x) = ax = exp( ln(a) x ) 9.1 ******** Notice that this tells us that 1 raised to any power x is equal to 1, 1x = exp(ln(1)x) = exp(0) = 1 9.2 This is obvious for x = n, and it is also true for our selected "interpolating function" we call ax. All of a sudden, we can now calculate things! For example, 3π = exp( ln(3)π). Since we know all about the functions ln and exp (imagine detailed tables for both), we know that 3π = exp( ln(3)π) = exp(1.0986..*π) = exp(3.451..) = 31.544.. And similarly, it we want the 7th root of 3, 31/7 = exp( ln(3)/7) = exp(1.0986../7) = exp(0.1569...) ≈ 1.17 1.177 = 3.0012 Now consider the derivative of 9.1. We get that ∂xax = ∂x exp( ln(a) x ) = ln(a) ∂xlna exp( ln(a) x ) = ln(a) exp( ln(a) x ) = ln(a) ax 9.3 We know there is a strange number e = 2.71.... which has the property ln(e) = 1. It is just the point x where ln(x) = 1, and we know such a point exists just looking at the graph for ln(x), We can compute e to any accuracy we want using something like Newton iteration. If we set a = e in the above, we get f(x) = ex = exp( ln(e) x ) = exp(x). // setting a = e Finally we arrive at a conclusion we have been seeking: Theorem: The function exp(x) which is the inverse of the ln(x) function can be represented as ex, which is to say, as a special real number e raised to the power x. Of course the "meaning of ex " is based on the meaning of the interpolating function ax, as developed above . ex is just a special case of ax. Remember that before we started this writeup, the function ax only had "meaning" for x = positive integers, and the same would be true for the special case ex. So if at the very start someone told us that exp(x) = ex, we would say that was not a very useful fact since we don't know what it means to raise a constant to a power which is not a positive integer. It then follows that f(x) = ax = exp( ln(a) x ) = e[xln(a)] This is our "interpolating function". That is, this can be regarded as the definition of ax for general x. We can if we want write this as e[xln(a)], but that conveys no new information because exp( ln(a) x ) is the definition of exlna. Below we will use both notations interchangeably. (10) We now want to show certain other properties. First off, ln(ax) = ln(eln(a)x) = ln(a) x. Thus: ln(ax) = x ln(a) or ln(xy) = y ln(x) 10.1 Next, start off with xy = xy, surely true. Apply ln to both sides. Write LHS = ln exp(xy) and RHS = y ln (exp(x)) = ln (exp(x)y) using 10.1 . Then apply exp to LHS = RHS to get exp(xy) = exp(x)y, [exp(x)]y = exp(xy) or (ex)y = exy 10.2 Now consider this application of 10.2: (ax)y = (exln(a))y = exyln(a) = axy , Thus, (ax)y = axy 10.3 Now let's throw in an application of 8.10 above: a-x = e-xlna = 1/ exlna = 1/ax so we have a-x = 1/ax 10.4 And here is one more using 8.9: axay = exlna eylna = e[xlna + ylna] = elna(x+y) = ax+y so we have axay = ax+y 10.5 Here I am trying to collect useful facts about our "interpolating function" f(x) = ax. (11) Finally, we want to verify that our various requirements we made concerning the interpolating function f(x) = ax are borne out in our explicit form for ax , namely, ax = exp(x ln(a)). Looking back at (6) above, we see that our first assumption was axay = ax+y . But this is the content of 10.5 above, so our explicit function does in fact satisfy this requirement we made on ax. The second thing we assumed was that a0 = 1, but a0 = exp(0 ln(a)) = exp(0) = 1, so that one is met too. In addition, we have generalized our result (6b) for n replaced by any real number, as shown in 10.3. And let's not forget the requirement that our ax be right at the integers! We have an = exp(n ln(a)) = [ exp(ln(a))]n = [a]n = a*a..a n times. (12) We can then go on to define more general functions like this: loga(y) ≡ ln(y)/ ln(a) 12.1 Then we get loga[ ax] ≡ ln [ax]/ ln(a) = x ln(a)/ln(a) = x So this gives us the association: y = ax x = loga(y) 12.2 The various ln properties then follow through. For example, loga(xy) = ln(xy)/ln(a) = ln(x)/ln(a) + ln(y)/ln(a) = loga(x) + loga(y). loga(1) = ln(1)/ ln(a) = 0 loga(1/x) = ln(1/x)/ln(a) = - ln(x)/ln(a) = - loga(x) and so on. Notice from 12.1 that loge(x) = ln(y) since ln(e) = 1. (13) Now we collect our various "formulas" in one place: ln(x) ≡ !Syntax Error, Idt/t just the area under this curve 1/x starting at 1 exp(x) ≡ ln-1(x) exp(ln(x)) = x ln(x) = exp-1(x) ln(exp(x)) = x ∂x ln(x) = 1/x 8.1 ln(1) = 0 8.2 ∂x(exp(x)) = exp(x) 8.3 "is its own derivative" exp(0) = 1 8.4 ln(xy) = ln(x) + ln(y) 8.5 ln(1/x) = - ln(x) 8.7 ln(x/y) = ln(x) – ln(y) 8.8 exp(x) exp(y) = exp(x+y) 8.9 // Ahlfors p 43 credit exp(-x) = 1/exp(x) 8.10 f(x) = ax = exp( ln(a) x ) 9.1 ******** 1x = exp(ln(1)x) = exp(0) = 1 9.2 f(x) = ax = exp( ln(a) x ) = e[xln(a)] ln(ax) = x ln(a) or ln(xy) = y ln(x) 10.1 [exp(x)]y = exp(xy) or (ex)y = exy 10.2 (ax)y = axy 10.3 a-x = 1/ax 10.4 axay = ax+y 10.5 loga(y) ≡ ln(y)/ ln(a) 12.1 y = ax x = loga(y) 12.2 loga(xy) = loga(x) + loga(y). loga(1) = = 0 loga(1/x) = - loga(x) loga(x/y) = loga(x) - loga(y). loge(x) = ln(y) (14) Concluding comments. We wanted to find a meaning for ax for general real x. At the start we only had a meaning when x = positive integer n. We imposed certain reasonable-sounding requirements on ax and we ended up with f(x) = ax = exp( ln(a) x ) as our "meaning". This function is a "continuation" of an to values of the exponent that are not positive integers (f(x) is an interpolation between the f(n)). [ I imagine some integral representation of ax which conveys this idea, a path not taken above. ] Our resulting formula for ax was shown to meet all the requirements we imposed on the function. We needed a lot of historical knowledge to arrive at our definition of ax. First, we had to know calculus, at least of one variable. So you cannot define ax for general x without calculus, you can't do it just with the concepts of simple algebra, whereas an could be handled in the simple algebra world. Within the world of calculus, we had to define the related functions ln(x) and exp(x), and both these functions are needed in our final expression that ax = exp(x ln(a)). [ As noted below, the world learned about calculus and these functions at roughly the same time, perhaps 1675 in the midst of Newton and Leibniz. Logarithms as a blind tool to do calculations were known earlier in 1614 (and to some extent before that). ] Once we reached our result for the interpolating formula ax = exp(x ln(a)), we could set a = e and then say that ex = exp(x). We realized that this is just a special case of ax and that we can use either notation we want for the exponential function. Once we found out about ax, we knew what it means to raise a number a to some arbitrary real exponent x. Then ex is just a special case of ax. The notation exp(x) indicates a certain function, but the notation ex really does say to raise a certain constant e to the power x using our ax interpolation formula as the meaning of doing this. It is amusing to ponder how different civilizations might arrive at these same conclusions. They must have a concept of ln(x), of exp(x) and of e. Perhaps they are not in a base-10 number system, but the concepts still survive. There is a nice web page on "the number e". (15) History of Logarithms. Before the age of electronic calculators and computers, people still needed to calculate things like the product or quotient of two numbers. Every grade school student knows (should know) how hard it is to manually compute 2.354 * 12.341. So-called long division is even harder to do, more painful. In the 1600's, as reviewed below, the basic idea of using log tables to do calculations was discovered. It is based on our formula above, for example loga(xy) = loga(x) + loga(y). loga(x/y) = loga(x) - loga(y). The idea is to pick some "base" a, and make one huge table showing loga(x) for a large range of x values. Certain tricks are available for a given base a such that not ALL x values need be in the table, just one "octave" of values, so to speak. Then when you wanted to multiply or divide two numbers, you looked up their logs in the table, added or subtracted these logs, then did a reverse lookup in the table to find the product. For example, suppose we want to know 2.354 * 12.341 and suppose we use base a = 10. Then we look up log10(2.354) = .3718 and log10(12.341) = 1.0914. We add the logs to get 1.463. We then scan our table to find the number whose log is 1.463 and we find that log10(29.050) = 1.463, and we conclude that 2.354 * 12.341 must be 29.050. We now refer to 29.050 as the anti-logarithm of 1.463. So you could say logax = y so that x = antilogay. Some wiki: The method of logarithms was first publicly propounded in 1614, in a book entitled Mirifici Logarithmorum Canonis Descriptio, by John Napier, Baron of Merchiston, in Scotland.[13] (Joost Bürgi independently discovered logarithms; however, he did not publish his discovery until four years after Napier.) Early resistance to the use of logarithms was muted by Kepler's enthusiastic support and his publication of a clear and impeccable explanation of how they worked.[14] At first, Napier called logarithms "artificial numbers" [ a term I might have used ] and antilogarithms "natural numbers". Later, Napier formed the word logarithm to mean a number that indicates a ratio: λόγος (logos) meaning proportion, and ἀριθμός (arithmos) meaning number. Napier chose that because the difference of two logarithms determines the ratio of the numbers they represent. So we have logs in 1614, before both Leibnitz and Newton were born. There was no "calculus" at this time. Logs were "discovered" as a method to assist in engineering and astronomy calculations in the Kepler era. A. Digression on Slide Rules. The famous "slide rule" uses two sliding sticks to add or subtract the logs of numbers. The sticks are labeled with numbers like 2.354 and 12.341, but these labels are located on the sticks at positions that are in fact the log10 of the numbers labeled. So distance along a stick is the log of the labeling number. You use two sticks as two rulers put "end to end" to add the distances. Here is an example: The bottom piece is the first stick (D scale), and the sliding middle part is the second stick (C scale). This stick is placed at the point marked "2" on the bottom stick, so the distance between 1 and 2 on the bottom stick is log 2 (in some distance units). If we then go a distance from "1" to "3" on the middle stick, we are adding log 2 + log 3. We look on the bottom stick and we see the label "6" at that point, so we know that log 2 + log 3 = log 6. Therefore, we know that 2*3 = 6, and so we have our answer 6. This same slide rule position would be used if we wanted to compute 2*2.5, which you can see is "5". In fact, this position gives 2 times any number you want. It seems that 2*π ≈ 6.3 on the far right. The same slide rule setting could be used to compute 6/2 = 3, so a slide rule can add or subtract "distances" = logs, and hence can do multiplication and division. Normally there is a clear plastic slider that helps you line things up. Ignore everything here but the slider. Other scales on the slide rule provided other functionality, such as computing powers and roots or computing trig functions. For example, we know that loga(x2) = 2 loga(x) so you could have two fixed scales, one where distance is loga(x) and the other where distance is 2 loga(x). But they are labeled with x and x2 values. These scales are called the D and A scales For example, if you look at 3 on the D scale above, you see 9 on the A scale, so 32 = 9. The middle sliding stick is of course not needed for such computations. And of course = 3. Each scale is only good for a specific power. With two fixed scales you can have a little table f(x) and we just looked at the case f(x) = x2. The K scale (not shown) does x3. Other scales did f(x) = tan(x), and of course the distance on the scale is really logatan(x) so you could compute then 5.34 * tan(31 degrees) or whatever. William Oughtred and others developed the slide rule in the 1600s based on the emerging work on logarithms by John Napier. Before the advent of the pocket calculator, it was the most commonly used calculation tool in science and engineering. The use of slide rules continued to grow through the 1950s and 1960s even as digital computing devices were being gradually introduced; but around 1974 the electronic scientific calculator made it largely obsolete and most suppliers exited the business. You can see looking at the above picture that you might located a number on the D scale 4.53. In general, you could expect at most 3 decimal places of "precision" on a slide rule like this. Wiki claims some Germans made a huge one: Typically the divisions mark a scale to a precision of two significant figures, and the user estimates the third figure. Some high-end slide rules have magnifying cursors that make the markings easier to see. Such cursors can effectively double the accuracy of readings, permitting a 10-inch slide rule to serve as well as a 20-inch. Astronomical work also required fine computations, and in 19th century Germany a steel slide rule about 2 meters long was used at one observatory. It had a microscope attached, giving it accuracy to six decimal places. B. Personal History. My high school years were 1962-1966, college 1967-1970. During most of this time I had at least one slide rule. I have only one left in my "memorabilia" box. I had a circular one at one time but it is lost. In later Harvard years, I saw Wang calculators in William James, people would reserve time on them. Then in Berkeley, the HP-35 arrived and I eventually bought one and I still have it and I think it still works. I recall an early price of $600, though, and only the "rich kids" at Berkeley had them. Another step toward the replacement of slide rules with electronics was the development of electronic calculators for scientific and engineering use. The first included the Wang Laboratories LOCI-2, [7] introduced in 1965, which used logarithms for multiplication and division and the Hewlett-Packard HP-9100, introduced in 1968. [8] The HP-9100 had trigonometric functions (sin, cos, tan) in addition to exponentials and logarithms. It used the CORDIC (coordinate rotation digital computer) algorithm, [9] which allows for calculation of trigonometric functions using only shift and add operations. This method facilitated the development of ever smaller scientific calculators. The era of the slide rule ended with the launch of pocket-sized scientific calculators, of which the 1972 Hewlett-Packard HP-35 was the first. Such calculators became known as "slide rule" calculators, since they could perform most, or all the functions of a slide rule. Introduced at US$395 [ $2035 2009 ] , even this was considered expensive for most students. But by 1975, basic four-function electronic calculators could be purchased for less than $50. By 1976 the TI-30 offered a scientific calculator for less than $25. After this time, the market for slide rules dwindled quickly as small scientific calculators became affordable. (16) History of ln(x) and exp(x) functions. Based on the snippets below, I would guess that the notions of exp(x) and ln(x) as discussed above first appeared around 1675 with Leibnitz, Newton and Bernoulli. These functions must have been present at the "founding of calculus". One must have thought right away about the problem of the integral of 1/x and the inverse of that function. It would be fair to say that Johann Bernoulli began the study of the calculus of the exponential function in 1697 when he published Principia calculi exponentialium seu percurrentium. The work involves the calculation of various exponential series and many results are achieved with term by term integration. So much of our mathematical notation is due to Euler that it will come as no surprise to find that the notation e for this number is due to him. The claim which has sometimes been made, however, that Euler used the letter e because it was the first letter of his name is ridiculous. It is probably not even the case that the e comes from "exponential", but it may have just be the next vowel after "a" and Euler was already using the notation "a" in his work. Whatever the reason, the notation e made its first appearance in a letter Euler wrote to Goldbach in 1731. He made various discoveries regarding e in the following years, but it was not until 1748 when Euler published Introductio in Analysin infinitorum that he gave a full treatment of the ideas surrounding e. He showed that e = 1 + 1/1! + 1/2! + 1/3! + ... and that e is the limit of (1 + 1/n)n as n tends to infinity. Euler gave an approximation for e to 18 decimal places, e = 2.718281828459045235 without saying where this came from. It is likely that he calculated the value himself, but if so there is no indication of how this was done. In fact taking about 20 terms of 1 + 1/1! + 1/2! + 1/3! + ... will give the accuracy which Euler gave. Among other interesting results in this work is the connection between the sine and cosine functions and the complex exponential function, which Euler deduced using De Moivre's formula. Leibniz is credited, along with Isaac Newton, with the discovery of infinitesimal calculus. According to Leibniz's notebooks, a critical breakthrough occurred on 11 November 1675, when he employed integral calculus for the first time to find the area under a function y = ƒ(x). He introduced several notations used to this day, for instance the integral sign ∫ representing an elongated S, from the Latin word summa and the d used for differentials, from the Latin word differentia. This ingenious and suggestive notation for the calculus is probably his most enduring mathematical legacy. Leibniz did not publish anything about his calculus until 1684.[22] The product rule of differential calculus is still called "Leibniz's law". In addition, the theorem that tells how and when to differentiate under the integral sign is called the Leibniz integral rule.