articles

From Quant Research to Parallel CPU and GPU Without Rewriting the Strategy

Table of contents

Python is a natural language for quantitative research. It is concise, interactive, and backed by an enormous numerical ecosystem. A pricing model or trading idea can often go from a notebook to a working experiment in a few dozen lines.

The problem comes when that experiment needs to scale.

Consider Monte Carlo simulation. The implementation is often straightforward: generate paths, evolve a stochastic process, calculate a payoff, and aggregate the results. But realistic workloads can require hundreds of thousands or millions of paths, each containing hundreds or thousands of time steps.

At that point, the obvious Python implementation can become painfully slow.

The usual answer is to rewrite the expensive part: vectorize it with NumPy, move the kernel into C++, introduce a JIT compiler, or implement a separate GPU version in CUDA. Each can work, but each also changes the way the program is written.

With Codon, we can take a different approach: in this post, we’ll start with an ordinary Python Monte Carlo simulation, then we’ll compile the same program with Codon, parallelize it across CPU cores, and finally move the simulation to the GPU. Importantly, across this whole process, the algorithm itself barely changes.

A Monte Carlo pricing problem

We’ll use an arithmetic-average Asian call option as a compact example.

Under geometric Brownian motion, each simulated price evolves according to

$$S_{t+\Delta t}
=
S_t
\exp\left[
\left(r-\frac{1}{2}\sigma^2\right)\Delta t
+
\sigma\sqrt{\Delta t}Z_t
\right],$$

where $Z_t$ is standard normal.

Unlike a vanilla European option, the payoff depends on the entire path:

$$\max\left(
\frac{1}{N}\sum_{t=1}^{N}S_t-K,\;0
\right).$$

That makes it a useful computational example: every path contains a sequential time evolution, while the paths themselves are independent and therefore naturally parallel.

Suppose we want to simulate 500,000 paths with 2,048 time steps each. That’s more than one billion price updates.

The core implementation is remarkably small:

for i in range(n_paths):
    s = S0
    s_sum = 0.0

    for t in range(n_steps):
        s *= np.exp(drift + diff * Z[t, i])
        s_sum += s

    avg = s_sum / n_steps
    payoffs[i] = max(avg - K, 0.0)

This is also the way many quants would naturally write the calculation: one outer loop over scenarios and one inner loop describing the path.

There’s just one problem: in CPython, executing a billion iterations is not especially practical.

Step 1: ordinary Python

Let's put the simulation into a complete program:

import time
import numpy as np

def var_cvar(x, alpha=0.99):
    x = np.sort(x)
    k = int(alpha * (len(x) - 1))
    return float(x[k]), float(x[k:].mean())

def main():
    n_paths = 500_000
    n_steps = 2048
    alpha = 0.99

    S0 = 100.0
    K = 100.0
    r = 0.03
    vol = 0.25
    T = 1.0

    dt = T / n_steps
    drift = (r - 0.5 * vol * vol) * dt
    diff = vol * np.sqrt(dt)

    rng = np.random.default_rng(123)
    Z = rng.standard_normal((n_steps, n_paths))

    payoffs = np.empty(n_paths)

    t0 = time.time()

    for i in range(n_paths):
        s = S0
        s_sum = 0.0

        for t in range(n_steps):
            s *= np.exp(drift + diff * Z[t, i])
            s_sum += s

        avg = s_sum / n_steps
        payoffs[i] = max(avg - K, 0.0)

    var, cvar = var_cvar(payoffs, alpha)

    print(f"VaR({alpha:.2f}) = {var:.6f}")
    print(f"CVaR({alpha:.2f}) = {cvar:.6f}")
    print(f"elapsed_sec = {time.time() - t0:.3f}")

if __name__ == "__main__":
    main()

Nothing fancy here; this is deliberately written as straightforward Python. On our benchmark system, the simulation takes:

CPython: 1150.6 seconds

That gives us our baseline.

Measurements were run on Ubuntu 22.04 (x86-64) with an Intel Xeon Platinum 8259CL at 2.50 GHz, with 2 physical cores / 4 hardware threads available and 15 GiB of system memory. GPU benchmarks used an NVIDIA Tesla T4 (compute capability 7.5) with 15 GiB of VRAM. All Codon benchmarks used Codon 0.20.1.

Step 2: compile with Codon

Codon is a high-performance Python compiler that compiles Python code to native machine code through LLVM.

For numerical code like this, that means the loops above don’t have to execute through the Python interpreter. Scalars can become native machine values, array operations can operate directly on native data, and the nested simulation loop becomes compiled code.

The important part is what we have to do to the source code: Nothing. The serial Codon version uses the exact same code as regular Python.

On the same machine:

CPython: 1150.6 seconds
Codon, serial: 24.7 seconds — 47x faster

This is already useful for research code containing loops that don’t map naturally onto a handful of NumPy operations. But Monte Carlo has another property we haven’t exploited: every path is independent.

Step 3: use all the CPU cores

The inner time loop is sequential: $S_{t+1}$ depends on $S_t$. The outer loop, however, is not, meaning the simulation is embarrassingly parallel across i. Codon supports multithreading via @par:

@par
for i in range(n_paths):

The compiler handles the details and machinery of parallelization, including loop outlining, assigning loop iterations to different threads, scheduling, and so on. The loop body stays exactly the same:

@par
for i in range(n_paths):
    s = S0
    s_sum = 0.0

    for t in range(n_steps):
        s *= np.exp(drift + diff * Z[t, i])
        s_sum += s

    avg = s_sum / n_steps
    payoffs[i] = max(avg - K, 0.0)

Codon distributes iterations of the outer loop across CPU threads. Each worker handles independent paths while the sequential evolution of each individual path stays intact.

Our result becomes:

Codon, serial: 24.7 seconds

‍Codon, @par: 10.6 seconds — 2.4x faster than serial, 109x faster than Python

At this point we’ve gone from interpreted Python to a multicore native implementation by adding one annotation, all without rewriting the simulation.

Step 4: run the same simulation on GPU

Monte Carlo also happens to fit GPUs extremely well. We have hundreds of thousands of independent paths performing the same computation.

Traditionally, moving this calculation to a GPU means introducing a second implementation. The CPU research code remains in Python while the performance-critical version becomes a CUDA kernel or is expressed through a GPU-specific framework.

Codon can instead compile parallel Python loops for the GPU.

Our CPU loop was:

@par
for i in range(n_paths):

For the GPU version, it becomes:

@par(gpu=True)
for i in range(n_paths):

The complete computational kernel is still:

@par(gpu=True)
for i in range(n_paths):
    s = S0
    s_sum = 0.0

    for t in range(n_steps):
        s *= np.exp(drift + diff * Z[t, i])
        s_sum += s

    avg = s_sum / n_steps
    payoffs[i] = max(avg - K, 0.0)

Codon compiles the parallel loop for execution on the GPU rather than the CPU.

On our benchmark GPU:

Codon, GPU: 3.8 seconds — 6.5x faster than serial, 303x faster than Python

The important takeaway isn’t just the final number. It’s the progression that got us there:

# Python / serial Codon
for i in range(n_paths):
    ...

# Parallel CPU
@par
for i in range(n_paths):
    ...

# GPU
@par(gpu=True)
for i in range(n_paths):
    ...

The simulation itself did not change.

One model, three execution targets

Here's our final performance progression:

The Python implementation can remain the specification of the model. With Codon, you don’t need one version for research, another version for CPU production, and another kernel for the GPU. The same loops, branches, mathematical expressions, and model logic can be compiled for very different hardware.

Now, our Asian option example is intentionally simple, but the computational pattern is everywhere in quantitative finance.

Monte Carlo engines often contain logic that becomes increasingly awkward to express as bulk array operations: path-dependent payoffs, barriers, early-exercise conditions, transaction costs, scenario-specific state, portfolio rules, stochastic volatility, multiple correlated assets, or custom risk calculations. Beyond that, the model rarely stays fixed.

A researcher might start with:

payoffs[i] = max(avg - K, 0.0)

then add a barrier:

if max_price > barrier:
    payoffs[i] = 0.0

then add another state variable, another process, or an entirely different payoff.

With ordinary Python code, those changes are easy to express. The challenge is keeping that flexibility when performance matters. Codon’s approach is to compile the program rather than requiring the programmer to reformulate the program around the compiler. That distinction becomes particularly valuable for quantitative research, where the code is not merely executed frequently, but is also changed frequently.

Monte Carlo simulation is one example, but the same pattern applies to many workloads in quantitative finance: scenario generation, risk calculations, backtesting, portfolio simulations, calibration, signal computation, and other compute-intensive research pipelines. Python is already an excellent language for expressing those models, but the question is really whether reaching the required performance should force you to stop writing Python.

With Codon, it doesn’t have to: starting from a normal serial implementation, we took the same Monte Carlo simulation from Python, to native execution, to parallel CPU execution, to a GPU kernel. The only source-level changes required for parallel execution were single-line annotations. The end result was 300x performance improvement over the original Python execution.

Keep your models in Python. Give your quants a faster path to production. Exaloop integrates Codon into your research platform for native CPU and GPU execution. Talk to Exaloop.
Talk to Exaloop about bringing native Python performance into your research platform.

Bringing native Python performance into the research platform

The bigger opportunity is not simply making one Python program faster. It is making high-performance Python a capability of the research platform itself.

In many quantitative organizations, there is still a boundary between research and production. Researchers work in Python because it lets them iterate quickly, but once a model becomes computationally demanding or needs to run in production, performance engineering becomes a separate step. Code is handed to another team, rewritten or reimplemented, validated against the research version, and then maintained alongside it. As models evolve, that boundary becomes an ongoing engineering cost.

Codon offers a different path. By integrating Codon into the research and production infrastructure, teams can give researchers a direct path from the Python they use to explore an idea to native CPU, multicore, and GPU execution. The model can remain Python as it moves through that lifecycle, reducing the need for a separate high-performance implementation and shortening the distance between a research change and its production deployment.

This is where Exaloop can help beyond the open-source compiler. We work with teams to integrate Codon into their existing platforms and workflows—from compilation and hardware targeting to deployment and production infrastructure—so high-performance execution becomes something researchers can use directly rather than something they have to hand off to a separate development effort.

If you want your quants to be able to take Python from research to high-performance production without waiting for a rewrite, talk to us about bringing Codon into your research platform.

Stay in the loop

Join our mailing list to stay updated on new features, product releases and announcements.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.