Spiral Logo

Latest posts

Fusing GPU Kernels with Dynamic Dispatch

Oct 7, 2026 · by Alexander Droste · 10 min read

We designed Vortex to provide the fastest possible decompression speeds on modern CPUs. As we showcased in a previous blog post, our Fastlanes implementation can decode 100 billion integers per second, fully utilizing the CPU's support for vectorized execution and instruction pipelining.
Because of our investments in performance, we've seen users from across the community adopt Vortex as the core storage layer for next-generation analytics and observability systems, primarily built on server-class CPUs.
However, we built Vortex to be the data format for the next decade, and GPU-native data systems are becoming an increasingly important part of that future. Whether it's LLMs or VLAs, pre-training or RL pipelines, GPUs are on their way to becoming the dominant consumers of open file formats.
Over the past year, we've been making improvements across Vortex's encoding, layout, and IO subsystems to unlock the full performance of modern GPU hardware, with its SIMT parallel execution model and it's multi-TB memory bandwidth.
This post is a deep dive into one of the hardest pieces of engineering that went into ultimately pushing Vortex to over 2TB/sec for decoding, an order of magnitude improvement that will unlock advanced data loading use-cases and create a baseline for GPU-accelerated analytics.

GPUs, briefly

There are excellent sources of information on GPU hardware and how it differs from traditional superscalar CPUs. We particularly recommend Modal's GPU Glossary for a technical deep dive on all things CUDA.
For those with a background in software performance on CPUs, there are a few key things to understand before we can get to the meat of the post.
The first: unlike CPUs, which use vectorization to process multiple elements of data in a single instruction, modern GPUs have large number of streaming multiprocessors executing identical code in parallel over disjoint sets of input.
For example, an NVIDIA GH200 has 132 streaming multiprocessors with 128 CUDA cores each, roughly 17,000 in total, and can keep 270,000 GPU threads in flight at once.
Second, GPUs have tremendous memory bandwidth, but less of it than you might have on a server-class CPU. The same GH200 Hopper has up to 144GB of high-bandwidth memory that can be accessed at speeds of over 4TB/sec, while its onboard CPU has LPDDR5X RAM that can be accessed at roughly 500GB/sec.
The memory bandwidth and SIMT execution model create enormous opportunity to achieve a potential order of magnitude improvement in decoding throughput, if we can figure out how to implement and schedule our decoding kernels to optimize our utilization of the GPU's unique resources.

Encoding trees

How do we actually use a GPU for decompression? In Vortex, data is compressed as a tree of nested lightweight encodings. Each node describes an encoding, whose child arrays may themselves use other encodings.
The following example illustrates an encoding tree using a dictionary as the outer encoding, nesting ALP, frame-of-reference, and bitpacking encodings.
Code Icon
Dict<f32>
├── codes: BitPacked<u32>
│   └── packed codes, bit_width = 6
└── values: ALP<f32>
    └── FoR<i32>
        └── BitPacked<i32>
            └── packed values, bit_width = 6
The benchmark we are going to use for this article uses a synthetic column of 100 million f32 values (400 MB), with 64 distinct values in the range of 0.10 to 0.73. Compression reduces the size from 400 MB to ~76 MB, a 5.3× compression ratio.

Chaining kernels

On the CPU, we can invoke each decoding routine with an ordinary function call. The equivalent GPU approach launches a separate CUDA kernel for each encoding, but each launch carries additional overhead relative to the CPU. The CUDA driver must prepare and submit commands to the GPU, whose hardware then schedules thread blocks onto its streaming multiprocessors. A typical launch costs roughly 5–15 µs of CPU and driver time, even for kernels that do very little work.
After moving the compressed data to the GPU, we can decompress it with the following five kernel launches:
Separate launches
Dictionary values and codes decoded in separate branchesPackedvaluesPackedcodesBit-unpackFrame Of RefALP decodeBit-unpackDictionarylookupf32output
For reference, this is the performance when decoding the same tree on the CPU:
Code Icon
# Benchmark on an Arm Neoverse-V2 CPU using 64 threads.
time:  12.31 ms    # Average duration to decompress 100M f32s.
thrpt: 30.3 GiB/s  # Average throughput to decompress 100M f32s.
If we benchmark the chain of kernels to decompress 100M f32s, we land on:
Code Icon
# Benchmark on a GH200. Excludes copying data to the GPU.
time:  437.50 us     # Average duration to decompress 100M f32s.
thrpt: 851.50 GiB/s  # Average throughput to decompress 100M f32s.
Even with this straightforward approach, Vortex already decompresses close to a terabyte per second. But we can go faster.
As noted earlier, launch overhead is a major cost when chaining GPU kernels. Vortex pays this cost for every batch, a contiguous range of rows that Vortex decodes together. While this overhead may be modest for shallow encoding trees, it becomes particularly costly for complex and deeply nested trees. Another issue is memory locality: within each batch, standalone kernels write intermediate results to GPU global memory, such that subsequent kernels can read them back.
Chaining KernelsThe bit_unpack kernel first loads packed values from global memory through L2 and L1 into registers. Cache hits can avoid an HBM read. It then stores values_for through L2. The for_decode kernel loads that buffer through the cache hierarchy into registers, then stores values_alp. The alp_decode kernel loads values_alp back into registers. Registers do not carry intermediate values directly from one kernel to the next. Dashed branches from L2 to HBM represent cache misses and writeback, not mandatory traffic for every access. Global stores bypass the L1 data cache in this schematic.Chaining KernelsRegistersL1 / TEXL2DevicememoryBitpackingFrame of ReferenceALPInput buffers / Intermediate results written back from L2loadstoreloadvalues_forstoreloadvalues_alp

Fusing kernels

One solution for bypassing the cost of chaining GPU kernels is fusing them. We could JIT-compile the entire decoding plan for an encoding tree into a single CUDA kernel. This allows the compiler to specialize the kernel for that tree, but introduces significant compilation latency the first time we encounter it. This works well when only a limited set of encoding combinations occurs and compiled kernels can be reused. In Vortex, however, each compressed chunk can have a different encoding tree. Specializing for each new tree can therefore introduce repeated compilation latency.
Instead, we use a dynamic dispatch kernel that can execute different encoding trees without compiling a new kernel for each tree.

Dynamic dispatch

The dynamic dispatch kernel contains precompiled decoding implementations for a subset of the encodings supported on the CPU. We launch this single kernel with a physical plan that specifies which implementations to run and in what order.
At runtime, the kernel selects operations such as bit-unpacking, frame-of-reference decoding, ALP decoding, and dictionary lookups, all within the same launch.
Dynamic Dispatch Kernel
The plan selects decoding routines inside a single kernelRuntime planWhich operations to run, and in what orderOne kernel launchPrecompiled ahead of timeSelect routine from planBit-unpackFrame-of-referenceALP decodeDictionarylookupContinue executing planNext operationPlan completeDecoded output

Plan execution

Since Vortex is written in Rust and GPU kernels are written in CUDA, we need a shared representation of the encoding tree. Vortex does this via a C ABI that defines a physical plan.
A plan is a sequence of stages executed within the dynamic dispatch kernel. Each stage starts with a source operation that reads or unpacks N encoded input elements into M values (N → M). The input and output cardinalities may differ because a packed word can contain several values, which bit-unpacking expands into separate output elements. Scalar operations then transform each value, producing one output per input (M → M).
Source and scalar operations
Two-stage dictionary decoding: an input stage supplies dictionary values to the output stage's scalar lookupPlanStage 1 · dictionary valuesStage 2 · dictionary codesSource operationN → MBit-unpackScalar operationM → MAdd referenceScalar operationM → MALP decodeSource operationN → MBit-unpackScalar operationM → MDictionary lookupDecoded dictionaryin shared memory
While constructing the dynamic dispatch plan, we track the amount of shared memory the kernel will use as scratch memory and pass that byte count when launching the kernel from the Rust side.
Keeping register usage low is particularly challenging for a dynamic kernel. The compiler determines the register count for the whole kernel at compile time, even though a given plan uses only some of its decoding paths. Supporting all these paths can increase the register count per thread. Each streaming multiprocessor (SM) has a fixed register budget, so higher register usage can reduce the number of warps it can keep resident.
To measure how many warps our kernel keeps resident, we use NVIDIA Nsight Compute (ncu), a profiler that reads GPU hardware performance counters:
Code Icon
Metric Name                                         Metric Value
-------------------------------------------------   -------------
sm__warps_active.avg.pct_of_peak_sustained_active   95.30%
Occupancy is the fraction of an SM's available warp slots occupied by resident warps. Our dynamic kernel achieves 95.3% for this plan, giving the schedulers a large pool of warps to draw from when hiding memory and instruction latency.

Performance

Does dynamic dispatch give us any performance gains over chaining standalone kernels?
Decompression throughput GiB/s · higher is better
Chained kernels851.50 GiB/s
Dynamic dispatch1704.8 GiB/s
GH200 · 100M f32 values · decoded output throughput · host-to-device copy excluded.
Yep, that's a 2x speedup!

cuDF integration

Vortex is already able to decompress data on the GPU and pass it to cuDF. The integration is available from Rust, C++, and Python and uses the Arrow C Device interface to share GPU buffers without copying them.
In Python, vortex_cuda.to_cudf converts each Vortex batch into a cuDF DataFrame. Here we use it to sum the amount column of a sales file one batch at a time:
Code Icon
Python
import vortex
import vortex_cuda

total = 0.0
for batch in vortex.open("sales.vortex").scan(["amount"]):
    df = vortex_cuda.to_cudf(batch)
    total += df["amount"].sum()
print(total)
cuDF also powers GPU acceleration for pandas, Polars, and Apache Spark. These integrations can build on this support to read and decompress Vortex data on the GPU.
Further, we're working to make Vortex a core file format in cuDF, starting with its NDS-H benchmarks.

Single kernel, twice the throughput

On the GH200, Vortex reaches 1,704.8 GiB/s of decoded output throughput, up from 851.50 GiB/s with chained kernels, by fusing the decoding steps into one precompiled GPU kernel. Dynamic dispatch reduces launch overhead and avoids global-memory round trips for intermediate results while supporting different encoding trees without recompilation.
Want to try it? Give it a spin with our Python cuDF integration.