News

NVIDIA Wants GPU Code to Care Less About the GPU

NVIDIA is bringing native Rust support to CUDA. But the bigger story may be CUDA Tile, a new approach designed to make GPU software less dependent on the details of individual GPU architectures.

Published September 17, 2026 · Flopper.io Research

NVIDIA has announced new support for writing GPU kernels directly in Rust, bringing the increasingly popular systems programming language deeper into the CUDA ecosystem.

The announcement introduces two different ways of programming NVIDIA GPUs with Rust: cuda-oxide and cutile-rs.

The first gives developers the kind of low-level control traditionally associated with CUDA programming.

The second takes a different approach.

With cutile-rs and NVIDIA’s wider CUDA Tile programming model, developers describe operations on blocks of data and leave more of the work of mapping those operations onto the GPU to the compiler.

That might sound like a small change in how GPU software is written.

It could be much more important than that.

GPUs keep changing

Look at how quickly NVIDIA’s data-centre GPU range has developed.

The A100 was followed by the H100. Then came the H200, followed by the Blackwell generation and GPUs such as the B200 and B300.

Across those generations, almost everything has changed.

Compute performance has increased. Memory capacities have grown. Memory bandwidth has increased. New lower-precision data formats have appeared. Tensor Core capabilities have changed.

These differences matter when developers try to get the highest possible performance from the hardware.

Traditional CUDA programming exposes many of the details of how work runs on the GPU.

Developers can think about individual threads, groups of threads, memory movement and how work should be divided across the GPU.

That control is powerful.

It also means software can contain decisions that are closely connected to how a particular generation of GPU works.

NVIDIA is now developing another option.

What is CUDA Tile?

CUDA Tile lets developers work at a higher level.

Instead of describing what every individual GPU thread should do, developers describe operations on chunks of data called tiles.

The compiler and runtime then decide how that work should be mapped onto the actual GPU.

NVIDIA introduced CUDA Tile with CUDA 13.1 and described it as a programming model for writing GPU kernels above the traditional SIMT model.

The important idea is simple:

The developer describes more of the problem. The compiler handles more of the hardware.

NVIDIA says CUDA Tile can handle details including thread scheduling, hardware mapping and resource allocation.

It can also make use of specialised NVIDIA hardware without requiring the application to target every part of that hardware directly.

The aim is to make code easier to move across different NVIDIA GPU architectures.

Rust makes NVIDIA’s direction clearer

NVIDIA’s new Rust support provides a useful example of the difference.

The company is developing two Rust projects.

cuda-oxide follows the traditional SIMT approach. Developers still have direct control over concepts such as threads and memory.

cutile-rs uses CUDA Tile.

Here, developers work with tiles of data and the compiler decides how those tiles should run on the GPU.

NVIDIA’s advice is particularly interesting.

When choosing between the two approaches, NVIDIA says developers should start with Tile and move to SIMT when they need the additional control.

That suggests Tile isn’t simply an experiment sitting alongside traditional CUDA.

NVIDIA sees higher-level GPU programming as an important part of CUDA’s future.

Why does this matter?

Modern GPU software has an awkward problem.

The hardware is developing extremely quickly.

Software has to keep up.

A kernel that has been carefully optimised around one architecture may not automatically make the best use of the next generation.

The industry therefore spends significant engineering effort making software run efficiently across different GPUs.

CUDA Tile moves some of that responsibility towards the compiler.

A developer can describe an operation using tiles. The compiler can then decide how those tiles should be executed on the architecture underneath.

In theory, that gives NVIDIA more freedom to change the hardware without requiring developers to encode as many hardware-specific decisions into their applications.

This doesn’t make the underlying GPU irrelevant.

Quite the opposite.

The hardware still matters

A B200 is not suddenly the same thing as an H100 because they can run the same Tile program.

The physical characteristics of the GPUs remain different.

Memory capacity still matters.

Memory bandwidth still matters.

Compute performance still matters.

Power consumption still matters.

Supported numerical formats still matter.

And, for anyone actually buying or renting compute, price and availability still matter.

CUDA Tile is instead about reducing how much software needs to understand the implementation details behind those differences.

That creates an interesting change in the relationship between software and hardware.

NVIDIA GPUs are becoming more different at the hardware level while NVIDIA is trying to make those differences easier for software to handle.

CUDA Tile isn’t just about Rust

The Rust announcement is only the latest part of a wider effort.

CUDA Tile first appeared with CUDA 13.1, initially with Python support.

NVIDIA expanded the model to C++ with CUDA 13.3.

It has also been working on using CUDA Tile IR as a backend for Triton, the popular language and compiler used for writing high-performance GPU kernels.

Rust is now another frontend.

That is important because it shows CUDA Tile becoming a layer that multiple programming languages can target.

Instead of every language needing to understand every detail of every NVIDIA GPU architecture, they can potentially target the same intermediate Tile model.

NVIDIA can then optimise how that work reaches the hardware.

Why Rust?

Rust itself is also becoming more important in AI infrastructure.

A growing amount of infrastructure software is being written in Rust because it combines low-level performance with strong memory-safety guarantees.

NVIDIA is already using Rust elsewhere in its software stack.

Its Nova Linux driver is written in Rust, while NVIDIA Dynamo uses Rust at its core.

Until now, developers could control CUDA workloads from Rust, but writing the GPU kernel itself could still mean moving into another language.

NVIDIA’s new work aims to close that gap.

With cuda-oxide, Rust kernels can be compiled directly to PTX.

With cutile-rs, developers can instead write Rust using the higher-level CUDA Tile model.

There is another benefit here too: safety.

GPU programs can have thousands of threads accessing the same memory at the same time. Mistakes involving memory ownership can produce bugs that are difficult to reproduce.

The two Rust projects use Rust’s type and ownership systems to catch some of these problems before the program runs.

Don’t expect CUDA to disappear

None of this means traditional CUDA programming is going away.

There will always be workloads where developers want direct control over the hardware.

NVIDIA itself makes this distinction clear.

Its recommendation is to use Tile first, then move down to SIMT when an application needs control over things such as threads and memory.

That gives developers a choice.

For many kernels, a higher-level description may be enough.

For highly specialised workloads, developers can still work much closer to the GPU.

It is similar to what has happened elsewhere in computing: higher-level tools handle the common case while lower-level interfaces remain available when developers need maximum control.

It is still early

There is an important warning around the Rust announcement.

Neither of NVIDIA’s Rust projects is currently production-ready.

cuda-oxide is described as early alpha.

cutile-rs is further ahead and has already been published as a Rust package, but NVIDIA still warns that coverage is incomplete and APIs are likely to change.

The wider CUDA Tile project is also developing quickly.

CUDA 13.1 introduced the model. CUDA 13.3 expanded it into C++. Rust support is now appearing alongside it.

This is therefore better understood as a direction of travel rather than the finished state of GPU programming.

The bigger picture

For years, one of the advantages of CUDA has been the amount of control it gives developers over NVIDIA GPUs.

NVIDIA isn’t removing that.

Instead, it is building another layer above it.

CUDA Tile asks developers to describe larger pieces of computation and gives the compiler more responsibility for turning that work into efficient GPU execution.

Rust is now becoming part of that system.

If the approach works, software written today could carry fewer assumptions about the exact NVIDIA GPU architecture it will eventually run on.

That becomes increasingly useful as NVIDIA releases new GPU architectures at a faster pace.

The A100, H100, H200, B200 and B300 can have very different specifications underneath.

NVIDIA’s goal is to make software care less about some of those differences.

For GPU users, however, those hardware differences aren’t going anywhere.

Someone still has to decide which GPU provides the right memory, bandwidth, compute performance, power efficiency and price for a workload.

The compiler may increasingly handle how software runs on the GPU.

Choosing which GPU to run it on remains a very different problem.

© 2026 Flopper.io - Compare the GPUs Powering AI