My day job is software for industrial X-ray computed tomography. An instrument takes thousands of X-ray projections of an object as it rotates, and the reconstruction software turns them into a 3D volume. This page describes a year of performance work on a GPU reconstruction engine and the tools around it.
The code belongs to my employer, so this page names no products, customers, data sets or file formats, and shows results as relative values measured on the same hardware before and after.
The problem#
The engine is a cone-beam reconstruction pipeline (filtered back-projection, FDK) written in C++ and CUDA: a chain of GPU operations from raw projections to a 3D volume, with a command-line front end, a gRPC server and Python hooks. The volumes it produces are routinely larger than the GPU’s memory and often larger than the host’s RAM, so the volume is reconstructed in slabs.
A profile of the starting point showed where the time went:
- the whole filtered projection set was uploaded to the GPU again for every slab: per run, the host-to-device traffic was about 35× the size of the input;
- almost half of the CUDA API time was spent in blocking stream synchronisation, and transfers overlapped with computation only about 4% of the time;
- the volume writer was slow and blocked the GPU while it wrote;
- smaller GPUs crashed with out-of-memory errors instead of adapting, and every run leaked host memory;
- artifact correction was silently skipped in some modes, and failed runs reported success.
What I changed#
- Asynchronous data movement. Hidden copies and blocking synchronisations were replaced by explicit asynchronous transfers ordered by CUDA events, on non-blocking streams.
- Caches instead of re-uploads. A pinned host cache with leases, a device memory pool, and a projection cache that stays on the GPU across slabs.
- Upload only what a slab needs. For each slab, the cone-beam geometry tells exactly which detector rows project onto it. Only those rows are uploaded, and only the difference from the previous slab. The result is bit-identical.
- Streaming artifact correction. Ring and streak removal works in small sinogram bands, in place, instead of on full-size copies of the data.
- A memory model you can trust. Host and GPU memory needs are computed before the run. A machine that cannot do the job gets a clear refusal with the numbers in under a second, instead of a crash an hour later.
- A parallel writer. Positional writes from a pinned ring buffer, with the number of writers chosen by the storage type (spinning disks get one), and crash-safe output that appears only when complete.
- Fewer ways to get it wrong. 18 subcommands became 3 modes, 16 environment switches became validated configuration, and thousands of lines of dead code went away.
- Errors must surface. A reliability audit fixed race conditions, 64-bit indexing for very large volumes and swallowed CUDA errors; static analysis went from over two thousand findings to zero.
Results#
All numbers compare the same data on the same machine, with medians of repeated, interleaved runs.
| What | Change |
|---|---|
| End-to-end reconstruction, equal work | more than 2× faster (−54% to −60% wall time) |
| Production mode with artifact correction | −44% to −50%, while doing more work than before |
| Back-projection stage | −57% to −61% |
| Data moved from host to GPU per run | −89% to −99.8% |
| Time in blocking stream synchronisation | −92% |
| Volume writer throughput | 2.3× higher; the GPU idle tail at the end of a run is about 75× shorter |
| Same run on a spinning disk | 3.3× faster |
| Projection reader | 4.1× faster (−20% for the whole run) |
| Run start-up (memory preparation) | 110× faster |
| Host memory for artifact correction | −63%, bit-identical output, −25% time |
| Interactive single-slice preview | up to 16.7× faster on the largest data |
| The largest data sets | from “does not fit in memory” to an ordinary run |
| Host memory leaks, GPU memory errors | none (checked with compute-sanitizer) |
| Unit tests | about 3× more |
Not every idea won, and the losers are part of the result. A non-atomic back-projection was 23% slower, register tiling gave nothing because the kernel is limited by texture fetches, and a hierarchical back-projection failed the quality, speed and memory gates. A CPU SIMD study sped individual loops up 19× but changed nothing end to end, because CPU work was under 1% of the critical path, so it was reported as a refuted hypothesis rather than shipped as a win.
Proving nothing changed#
Speed is worthless if the image changes. Every change had to pass a gate:
- Bit-exact golden volumes: the output of each reference reconstruction is hashed, and a change must reproduce it exactly (multi-worker runs may differ by one least significant bit).
- Analytic phantoms with exactly computed projections as ground truth, next to real scans.
- “Red proofs”: a new test must fail on a build with the fix removed.
- This process found a subtle defect in the original code: a detector-row drift that grew from slab to slab. After the fix, a volume made in one slab and the same volume made in eight are bit-identical.
Seeing inside the pipeline#
A GPU pipeline is a black box: data goes in, a volume comes out. I built a live pipeline viewer to open it:
- operations in the pipeline publish downsampled copies of their GPU data into a lock-free shared-memory ring; a separate viewer process reads them without copying, so a viewer crash can never stop a reconstruction, and the reconstruction never waits for the viewer;
- an optional blocking tap turns it into a step debugger for the GPU pipeline, with a frame timeline and a live graph of the pipeline;
- the viewer also controls the reconstruction server over gRPC: mode, options, live progress;
- offline viewers for volumes and projections, with orthogonal slices and a ray-cast 3D view that shows an instant preview and then streams in full resolution;
- an A/B comparison of two volumes that reads only the slices on screen, so two volumes of hundreds of gigabytes can be compared over a network share.
A harness that proves images are the same#
Changes to reconstruction need proof at real scale, not on a toy example. A Python verification harness runs tiers of scenarios (real scans and analytic phantoms) against pinned baselines and compares every slice with several image metrics (SSIM, perceptual hashing, an anti-aliasing-aware pixel difference and NVIDIA FLIP). It writes HTML reports and a history of results, and it is careful with disks that hold terabytes: it checks free space first and moves baselines instead of copying them.
The charting library#
The same year I worked on a scientific charting library for an instrument desktop suite: a WPF control rendered with DirectX.
- Migration from .NET Framework with Direct3D 10 to .NET 9 with Direct3D 11: one shared GPU device for all charts instead of one per chart, RAII wrappers for GPU resources, and a render pump that coalesces updates.
- Deferred surfaces: with 100 charts open, about 12× faster and 8.6× less memory. The new design aims at one render per settings change, where the old control did five to ten.
- Visual regression CI: the old and the new renderer both draw the same scenarios off-screen, and four image metrics vote on every pair, which tolerates rasteriser differences between Direct3D versions but catches real regressions. A benchmark gate stops a merge that makes anything slower, and a leak gate opens and closes hundreds of charts and expects nothing left behind.
- A target architecture for the next step: one chart model with views on top, rendering surfaces from a pool, a single scheduler, and a frozen public API, planned in tracks with a weekly benchmark gate.
Tools of the trade#
CUDA (texture-based back-projection, cuFFT, half-precision storage, Nsight Systems and Nsight Compute, compute-sanitizer), C++20, out-of-core memory management, asynchronous streaming, AVX2 and OpenMP, gRPC, shared-memory IPC, Dear ImGui and OpenGL, Python with CuPy, WPF and Direct3D 11, static analysis (clang-tidy, MSVC, PVS-Studio), and an AI-assisted workflow with Claude Code: issue boards, isolated worktrees and verification gates.