Category: AI Software Engineering GPU
Difficulty: Medium

#127 CUDA Indexing Explained: Threads, Blocks and Grids You Can Play With đź§©

CUDA indexing is the part of GPU programming that confuses almost everyone at the start. Below you will find a school analogy, the one formula you need, and two interactive visualizations so you can see how threads, blocks and the grid fit together.

The problem: 1,000 apples, 1,000 helpers

Imagine you have 1,000 apples to wash and you hire 1,000 helpers. You give all of them the same instruction: “wash an apple”. If that is all you say, every helper grabs apple number 1.

So every helper needs to answer one question: which apple is mine?

That is all CUDA indexing is. A GPU kernel is one piece of code that thousands of threads run at the same time, and each thread uses its own index to pick its own piece of data.

The cast: a school

CUDA term School analogy Built-in variable
Thread one kid threadIdx (seat number inside the classroom)
Block one classroom blockIdx (classroom number inside the school)
Grid the whole school gridDim (how many classrooms)
Threads per block kids per classroom blockDim

The catch: seat numbers restart at 0 in every classroom. Seat 2 exists in classroom 0, classroom 1 and classroom 2. If every kid used only their seat number, three kids would grab the same apple.

The fix is a school-wide ID:

ID = (classroom number Ă— kids per classroom) + seat number

In CUDA code that is the famous line:

int i = blockIdx.x * blockDim.x + threadIdx.x;

Visualization 1: 1D indexing

Each colored box is a block, each small square is a thread. The top number of a square is threadIdx.x (it restarts in every block), the bottom number is the global index i. The row at the bottom is your data array.

How to play:

  • Hover (or tap) a thread to see the formula and which array element it works on.
  • Move numBlocks down until some array cells turn dashed red. Those elements have no thread and will silently never be processed.
  • Move n below the thread count. The extra threads fade out. They are idle because of the bounds check.
4
3
10

Grid: each box is a block, each square is a thread (top = threadIdx.x, bottom = global i)

What you just saw

  • threadIdx.x is local: it restarts at 0 in every block.
  • The global index i is unique: it keeps counting across blocks.
  • Every block has the same size, so the last block is often not completely needed. That is where idle threads come from.

What if the sizes do not match?

Say you have 45 apples but launch 5 blocks Ă— 5 threads = 25 threads. Threads get IDs 0 to 24, so apples 25 to 44 are never touched. There is no error message, your program just returns a half-finished result. In the visualization above, those are the dashed red cells.

Fix 1: launch enough threads

Round up when computing the number of blocks:

int threads = 5;
int blocks  = (n + threads - 1) / threads;   // ceil(45 / 5) = 9
kernel<<<blocks, threads>>>(a, b, out, n);

If it does not divide evenly (say 23 apples with 5 threads per block gives 5 blocks = 25 threads), the last two threads have no apple. Those two must do nothing, which is the job of the bounds check:

__global__ void add(const float* a, const float* b, float* out, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < n) {                 // threads with i >= n sit out
        out[i] = a[i] + b[i];
    }
}

Without if (i < n) the extra threads read and write memory outside your array. That can corrupt other data or crash, often silently. Idle threads are cheap, out-of-bounds access is not.

Fix 2: grid-stride loop

If you cannot (or do not want to) launch one thread per element, let each thread take several elements. After finishing one, it jumps ahead by the total number of threads:

__global__ void add(const float* a, const float* b, float* out, int n) {
    int i      = blockIdx.x * blockDim.x + threadIdx.x;
    int stride = gridDim.x * blockDim.x;       // total threads in the grid
    for (; i < n; i += stride) {
        out[i] = a[i] + b[i];
    }
}

With 25 threads and 45 elements, thread 0 handles elements 0 and 25, thread 1 handles 1 and 26, and so on. Every element is processed exactly once, and the kernel works no matter how many threads you launch.

Visualization 2: 2D indexing (images and matrices)

For an image, one number is not enough: each thread needs a column and a row. It is the same formula, applied twice:

int col = blockIdx.x * blockDim.x + threadIdx.x;
int row = blockIdx.y * blockDim.y + threadIdx.y;
int idx = row * width + col;       // plain row-major array indexing

Here the image is 8 pixels wide and 6 tall, split into 4Ă—3 thread blocks, so the grid is 2Ă—2 blocks. Each pixel shows its memory index and, in small text, (threadIdx.x, threadIdx.y). Hover over any pixel.

Image is 8 wide x 6 tall, blockDim = (4, 3), gridDim = (2, 2). Each pixel shows its memory index (row * width + col).

What to look for

  • Neighbors in x are neighbors in memory. Hover along a row and the index goes 8, 9, 10, 11 (+1 per step). Hover down a column and it jumps by 8, the image width. That is why threadIdx.x should map to the column: consecutive threads then read consecutive addresses (coalesced access), which is much faster.
  • The CUDA part is only col and row. row * width + col is ordinary row-major array indexing.
  • threadIdx is local. Pixel 6 in block (1,0) has threadIdx = (2,1) but global column 6. Block (0,0) also has a thread with threadIdx.x = 2, and it works on a different pixel.

Threads are jobs, cores are hands

A common question: is the number of threads the number of (tensor) cores? No.

  • Threads are the jobs you ask for in <<<blocks, threads>>>. You can launch millions.
  • Cores are the actual hardware, fixed by the chip. If you launch more threads than cores, the extra threads simply wait their turn.

Example with an NVIDIA T4:

Hardware part What it is T4 count
SM (streaming multiprocessor) The “classroom building” that your blocks are assigned to 40
CUDA cores Regular cores, one scalar operation at a time 2,560 (64 per SM)
Tensor cores Special units that multiply small matrices in one step 320 (8 per SM)

Inside an SM, threads run in groups of 32 called warps. Tensor cores do not change the indexing formula. You normally use them through libraries (cuBLAS, PyTorch) rather than through plain per-thread code.

Cheat sheet

Question Answer
Which element am I? i = blockIdx.x * blockDim.x + threadIdx.x
2D version col from x, row from y, then row * width + col
How many blocks do I need? (n + threads - 1) / threads
Why if (i < n)? The last block usually has extra threads that must sit out
More data than threads? Grid-stride loop: i += gridDim.x * blockDim.x
threadIdx vs i? threadIdx restarts in every block, i is globally unique

Ask yourself two questions in every kernel: “Which element am I?” and “Is that element actually valid?” Everything else is arithmetic.

Written on October 9, 2026