decimo

Benchmarks

2026-09-02 · main 92a323b · arm64, macOS-26.6.2-arm64-arm-64bit-Mach-O

Mojo 1.0.0 (ed45d567) · CPython 3.14.6 · libmpdec 4.0.1 and GMP 6.3.0, both linked from C

Each figure is the minimum over several rounds, on an idle machine. Regenerate with pixi run benchdoc.

BigDecimal

Operands of digits decimal digits at working precision, so that precision grows with the operands. libmpdec is the C library behind CPython’s decimal, timed without the interpreter. round drops the low half of the operand; parse builds the value from its string.

Verified this run: decimo and libmpdec produce the same 2000-digit product from the same 1000-digit operands.

Digits / precision Operation decimo libmpdec   CPython decimal  
9 / 28 add 10.6 ns 33.8 ns 3.20× faster 46.9 ns 4.43× faster
9 / 28 subtract 11.3 ns 29.7 ns 2.62× faster 47.1 ns 4.16× faster
9 / 28 multiply 13.0 ns 26.8 ns 2.06× faster 44.4 ns 3.41× faster
9 / 28 divide 57.0 ns 49.0 ns 1.16× slower 67.9 ns 1.19× faster
9 / 28 round 8.0 ns 34.6 ns 4.34× faster 69.6 ns 8.72× faster
9 / 28 parse 49.2 ns 39.9 ns 1.23× slower 83.6 ns 1.70× faster
1,000 / 1,000 add 46.0 ns 115.3 ns 2.51× faster 109.5 ns 2.38× faster
1,000 / 1,000 subtract 52.1 ns 93.8 ns 1.80× faster 85.1 ns 1.63× faster
1,000 / 1,000 multiply 1.04 µs 9.01 µs 8.66× faster 9.78 µs 9.40× faster
1,000 / 1,000 divide 5.72 µs 14.44 µs 2.53× faster 14.88 µs 2.60× faster
1,000 / 1,000 round 50.5 ns 95.8 ns 1.90× faster 120.6 ns 2.39× faster
1,000 / 1,000 parse 1.24 µs 1.64 µs 1.32× faster 1.85 µs 1.49× faster
100,000 / 100,000 add 2.25 µs 6.00 µs 2.67× faster 5.97 µs 2.65× faster
100,000 / 100,000 subtract 2.75 µs 3.70 µs 1.35× faster 3.59 µs 1.31× faster
100,000 / 100,000 multiply 1.87 ms 2.83 ms 1.52× faster 3.00 ms 1.61× faster
100,000 / 100,000 divide 4.70 ms 13.79 ms 2.93× faster 14.53 ms 3.09× faster
100,000 / 100,000 round 2.35 µs 4.50 µs 1.91× faster 4.41 µs 1.88× faster
100,000 / 100,000 parse 106.40 µs 159.50 µs 1.50× faster 206.15 µs 1.94× faster
1,000,000 / 1,000,000 add 24.50 µs 58.50 µs 2.39× faster 61.06 µs 2.49× faster
1,000,000 / 1,000,000 subtract 27.50 µs 35.50 µs 1.29× faster 36.29 µs 1.32× faster
1,000,000 / 1,000,000 multiply 25.15 ms 29.72 ms 1.18× faster 31.65 ms 1.26× faster
1,000,000 / 1,000,000 divide 124.35 ms 157.24 ms 1.26× faster 166.62 ms 1.34× faster
1,000,000 / 1,000,000 round 28.50 µs 42.50 µs 1.49× faster 41.06 µs 1.44× faster
1,000,000 / 1,000,000 parse 1.17 ms 1.58 ms 1.35× faster 1.85 ms 1.59× faster

sqrt, exp, ln, power

A fixed set, chosen before the results were seen. 2.3456789, and ** 1.5 for power. 100 000 digits is left out because mpd_exp needs 82 seconds there and mpd_ln 239 seconds.

Precision Operation decimo libmpdec  
28 sqrt 379.3 ns 889.2 ns 2.34× faster
28 exp 2.18 µs 4.59 µs 2.11× faster
28 ln 2.68 µs 12.83 µs 4.79× faster
28 power 6.75 µs 35.71 µs 5.29× faster
100 sqrt 2.09 µs 4.10 µs 1.96× faster
100 exp 6.05 µs 14.86 µs 2.46× faster
100 ln 13.55 µs 38.98 µs 2.88× faster
100 power 22.96 µs 80.38 µs 3.50× faster
1,000 sqrt 20.27 µs 215.89 µs 10.65× faster
1,000 exp 173.63 µs 4.46 ms 25.69× faster
1,000 ln 1.18 ms 5.54 ms 4.70× faster
1,000 power 1.42 ms 11.18 ms 7.89× faster
10,000 sqrt 569.00 µs 22.76 ms 40.01× faster
10,000 exp 24.41 ms 1.370 s 56.10× faster
10,000 ln 484.22 ms 3.224 s 6.66× faster
10,000 power 507.12 ms 4.636 s 9.14× faster

BigInt against GMP and CPython’s int

GMP is the reference implementation for big integers and the harder opponent, timed in C. CPython’s integers can only be reached through the interpreter, so those include its overhead — read the large sizes as the real ones. Division is 2n-by-n; two operands of the same width would give a one-digit quotient and measure nothing.

An mpz_t always lives on the heap. Adding two 100-digit values costs GMP 14.9 ns into a fresh result against 3.2 ns into a reused one, so most of it is malloc and free. decimo keeps small values inside the struct, which is why it leads at the small end and trails at the large one, where the algorithm is all that is left.

Digits Operation decimo GMP   CPython int  
10 add 5.2 ns 17.7 ns 3.41× faster 23.5 ns 4.53× faster
10 multiply 7.3 ns 15.5 ns 2.14× faster 29.8 ns 4.11× faster
10 floor divide 10.4 ns 19.1 ns 1.83× faster 51.1 ns 4.90× faster
10 sqrt 2.5 ns 15.2 ns 6.15× faster 30.0 ns 12.10× faster
100 add 7.0 ns 14.9 ns 2.14× faster 28.1 ns 4.04× faster
100 multiply 40.5 ns 26.2 ns 1.55× slower 112.4 ns 2.78× faster
100 floor divide 112.7 ns 71.5 ns 1.57× slower 302.8 ns 2.69× faster
100 sqrt 115.7 ns 57.4 ns 2.02× slower 314.1 ns 2.71× faster
1,000 add 33.8 ns 26.0 ns 1.30× slower 89.3 ns 2.64× faster
1,000 multiply 1.03 µs 597.8 ns 1.72× slower 5.55 µs 5.40× faster
1,000 floor divide 1.78 µs 1.20 µs 1.48× slower 13.26 µs 7.47× faster
1,000 sqrt 1.53 µs 354.0 ns 4.33× slower 4.46 µs 2.91× faster
10,000 add 275.0 ns 145.0 ns 1.90× slower 783.1 ns 2.85× faster
10,000 multiply 46.56 µs 18.71 µs 2.49× slower 270.60 µs 5.81× faster
10,000 floor divide 94.41 µs 50.16 µs 1.88× slower 552.64 µs 5.85× faster
10,000 sqrt 23.23 µs 7.73 µs 3.01× slower 268.85 µs 11.57× faster
100,000 add 2.75 µs 1.45 µs 1.90× slower 7.83 µs 2.85× faster
100,000 multiply 1.20 ms 429.15 µs 2.81× slower 8.83 ms 7.34× faster
100,000 floor divide 3.74 ms 1.12 ms 3.35× slower 18.48 ms 4.94× faster
100,000 sqrt 987.40 µs 291.95 µs 3.38× slower 8.70 ms 8.81× faster
1,000,000 add 25.50 µs 11.50 µs 2.22× slower 87.58 µs 3.43× faster
1,000,000 multiply 24.66 ms 5.97 ms 4.13× slower 375.64 ms 15.23× faster
1,000,000 floor divide 105.06 ms 16.19 ms 6.49× slower 755.80 ms 7.19× faster
1,000,000 sqrt 31.82 ms 5.69 ms 5.60× slower 355.04 ms 11.16× faster

pi()

Each measurement in a fresh process, since mpmath and MPFR cache the constant. Includes conversion to a decimal string, for every library.

Digits decimo mpmath + GMP MPFR mpmath (pure Python)
100 2.17 µs 48.67 µs 49.58 µs 32.58 µs
1,000 21.64 µs 114.71 µs 69.33 µs 174.75 µs
10,000 727.70 µs 882.92 µs 718.62 µs 6.08 ms
100,000 26.65 ms 15.70 ms 23.69 ms 186.04 ms
1,000,000 714.22 ms 279.97 ms 429.02 ms 8.709 s