Latest posts
Fusing GPU Kernels with Dynamic Dispatch
Oct 7, 2026 · by Alexander Droste · 10 min read
We designed Vortex to provide the
fastest possible decompression speeds on modern CPUs. As we showcased in a
previous blog post, our Fastlanes implementation can decode
100 billion integers per second,
fully utilizing the CPU's support for vectorized execution and instruction
pipelining.
Because of our investments in performance, we've seen users from across the
community adopt Vortex as the core storage layer for next-generation analytics
and observability systems, primarily built on server-class CPUs.
However, we built Vortex to be the data format for the next decade, and
GPU-native data systems are becoming an increasingly important part of that
future. Whether it's LLMs or VLAs, pre-training or RL pipelines, GPUs are on
their way to becoming the dominant consumers of open file formats.
Over the past year, we've been making improvements across Vortex's encoding,
layout, and IO subsystems to unlock the full performance of modern GPU hardware,
with its SIMT parallel execution model and it's multi-TB memory bandwidth.
This post is a deep dive into one of the hardest pieces of engineering that went
into ultimately pushing Vortex to over 2TB/sec for decoding, an order of
magnitude improvement that will unlock advanced data loading use-cases and
create a baseline for GPU-accelerated analytics.
GPUs, briefly
There are excellent sources of information on GPU hardware and how it differs
from traditional superscalar CPUs. We particularly recommend
Modal's GPU Glossary for a technical deep dive
on all things CUDA.
For those with a background in software performance on CPUs, there are a few key
things to understand before we can get to the meat of the post.
The first: unlike CPUs, which use vectorization to process multiple elements
of data in a single instruction, modern GPUs have large number of streaming
multiprocessors executing identical code in parallel over disjoint sets of
input.
For example, an NVIDIA GH200 has 132 streaming multiprocessors with 128 CUDA
cores each, roughly 17,000 in total, and can keep 270,000
GPU threads in flight
at once.
Second, GPUs have tremendous memory bandwidth, but less of it than you might
have on a server-class CPU. The same GH200 Hopper has up to 144GB of
high-bandwidth memory that can be accessed at speeds of over 4TB/sec, while
its onboard CPU has LPDDR5X RAM that can be accessed at roughly 500GB/sec.
The memory bandwidth and SIMT execution model create enormous opportunity to
achieve a potential order of magnitude improvement in decoding throughput, if we
can figure out how to implement and schedule our decoding kernels to optimize
our utilization of the GPU's unique resources.
Encoding trees
How do we actually use a GPU for decompression? In Vortex, data is compressed as
a
tree of nested lightweight encodings.
Each node describes an
encoding, whose
child arrays may themselves use other encodings.
The following example illustrates an encoding tree using a
dictionary as the
outer encoding, nesting ALP,
frame-of-reference, and
bitpacking
encodings.
Dict<f32>
├── codes: BitPacked<u32>
│ └── packed codes, bit_width = 6
└── values: ALP<f32>
└── FoR<i32>
└── BitPacked<i32>
└── packed values, bit_width = 6
The benchmark we are going to use for this article uses a synthetic column of
100 million
f32 values (400 MB), with 64 distinct values in the range of
0.10 to 0.73. Compression reduces the size from 400 MB to ~76 MB, a 5.3×
compression ratio.Chaining kernels
On the CPU, we can invoke each decoding routine with an ordinary function call.
The equivalent GPU approach launches a separate
CUDA kernel for each
encoding, but each launch carries additional overhead relative to the CPU. The
CUDA driver must
prepare and
submit commands
to the GPU, whose hardware then schedules
thread blocks
onto its
streaming multiprocessors.
A typical launch costs roughly 5–15 µs of CPU and driver time, even for kernels
that do very little work.
After moving the compressed data to the GPU, we can decompress it with the
following five kernel launches:
Separate launches
For reference, this is the performance when decoding the same tree on the CPU:
# Benchmark on an Arm Neoverse-V2 CPU using 64 threads.
time: 12.31 ms # Average duration to decompress 100M f32s.
thrpt: 30.3 GiB/s # Average throughput to decompress 100M f32s.
If we benchmark the chain of kernels to decompress 100M
f32s, we land on:# Benchmark on a GH200. Excludes copying data to the GPU.
time: 437.50 us # Average duration to decompress 100M f32s.
thrpt: 851.50 GiB/s # Average throughput to decompress 100M f32s.
Even with this straightforward approach, Vortex already decompresses close to a
terabyte per second. But we can go faster.
As noted earlier, launch overhead is a major cost when chaining GPU kernels.
Vortex pays this cost for every batch, a contiguous range of rows that Vortex
decodes together. While this overhead may be modest for shallow encoding trees,
it becomes particularly costly for complex and deeply nested trees. Another
issue is memory locality: within each batch, standalone kernels write
intermediate results to GPU global memory, such that subsequent kernels can read
them back.
Fusing kernels
One solution for bypassing the cost of chaining GPU kernels is fusing them. We
could JIT-compile the entire decoding plan for an encoding tree into a single
CUDA kernel. This allows the compiler to specialize the kernel for that tree,
but introduces significant compilation latency the first time we encounter it.
This works well when only a limited set of encoding combinations occurs and
compiled kernels can be reused. In Vortex, however, each compressed chunk can
have a different encoding tree. Specializing for each new tree can therefore
introduce repeated compilation latency.
Instead, we use a dynamic dispatch kernel that can execute different encoding
trees without compiling a new kernel for each tree.
Dynamic dispatch
The dynamic dispatch kernel contains precompiled decoding implementations for a
subset of the encodings supported on the CPU. We launch this single kernel with
a physical plan that specifies which implementations to run and in what order.
At runtime, the kernel
selects operations such as bit-unpacking, frame-of-reference decoding, ALP
decoding, and dictionary lookups, all within the same launch.
Dynamic Dispatch Kernel
Plan execution
Since Vortex is written in Rust and GPU kernels are written in CUDA, we need a
shared representation of the encoding tree. Vortex does this via a C ABI that
defines a physical plan.
A plan is a sequence of stages executed within the dynamic dispatch kernel. Each
stage starts with a source operation that reads or unpacks
N encoded input
elements into M values (N → M). The input and output cardinalities may
differ because a packed word can contain several values, which bit-unpacking
expands into separate output elements. Scalar operations then transform each
value, producing one output per input (M → M).Source and scalar operations
While constructing the dynamic dispatch plan, we track the amount of
shared memory
the kernel will use as scratch memory and pass that byte count when launching
the kernel from the Rust side.
Keeping register usage low is particularly challenging for a dynamic kernel. The
compiler determines the register count for the whole kernel at compile time,
even though a given plan uses only some of its decoding paths. Supporting all
these paths can increase the register count per thread. Each
streaming multiprocessor
(SM) has a fixed register budget, so higher register usage can reduce the number
of warps it can keep
resident.
To measure how many warps our kernel keeps resident, we use
NVIDIA Nsight Compute (
ncu),
a profiler that reads GPU hardware performance counters:Metric Name Metric Value
------------------------------------------------- -------------
sm__warps_active.avg.pct_of_peak_sustained_active 95.30%
Occupancy is the fraction of an
SM's available warp slots occupied by resident warps. Our dynamic kernel
achieves 95.3% for this plan, giving the schedulers a large pool of warps to
draw from when hiding memory and instruction latency.
Performance
Does dynamic dispatch give us any performance gains over chaining standalone
kernels?
Decompression throughput GiB/s · higher is better
Chained kernels851.50 GiB/s
Dynamic dispatch1704.8 GiB/s
Yep, that's a 2x speedup!
cuDF integration
Vortex is already able to decompress data on the GPU and pass it to cuDF. The
integration is available from
Rust,
C++, and
Python
and uses the Arrow C Device interface to share GPU buffers without copying them.
In Python,
vortex_cuda.to_cudf converts each Vortex batch into a cuDF
DataFrame. Here we use it to sum the amount column of a sales file one batch
at a time:Python
import vortex
import vortex_cuda
total = 0.0
for batch in vortex.open("sales.vortex").scan(["amount"]):
df = vortex_cuda.to_cudf(batch)
total += df["amount"].sum()
print(total)
cuDF also powers GPU acceleration for
pandas, Polars, and Apache Spark. These
integrations can build on this support to read and decompress Vortex data on the
GPU.
Further, we're working to
make Vortex a core file format in cuDF,
starting with its NDS-H benchmarks.
Single kernel, twice the throughput
On the GH200, Vortex reaches 1,704.8 GiB/s of decoded output throughput, up from
851.50 GiB/s with chained kernels, by fusing the decoding steps into one
precompiled GPU kernel. Dynamic dispatch reduces launch overhead and avoids
global-memory round trips for intermediate results while supporting different
encoding trees without recompilation.
Want to try it? Give it a spin with our
Python cuDF integration.