All posts

10 min read

Eight Numbers at a Time

One instruction, a whole row of data: an introduction to SIMD

If you open a photo on your phone and slide the brightness control, the picture keeps up with your thumb. There are about twelve million pixels in that photo, and every one of them has to be recalculated for every position of the slider.

01
The short version

When scaling clock speeds stalled around 2005, chip architectures shifted toward wider data paths. SIMD applies a single instruction to an entire vector of values simultaneously.

Think of

the single-file queue

Thirty-six million cars moving through a single toll booth. You can speed up the barrier gate up to a physical limit, but eventually heat and mechanical strain prevent going any faster.

The problem with one-by-one

A pixel is three numbers rather than a single dot of colour: how much red, how much green, how much blue, each running from 0 to 255. A twelve‑megapixel image therefore holds roughly thirty‑six million numbers in a row. The simplest way to brighten it is a loop that fetches a number, adds a constant, writes it back, and moves on, repeating thirty‑six million times for every position of the slider.

For decades, chips handled this by increasing clock speeds, scaling from megahertz into gigahertz. Around 2005, standard silicon scaling hit thermal and power limits (the power wall), and base frequencies leveled off in the low gigahertz range.

Architectures adapted by increasing per-cycle throughput instead of raw clock cycles. SIMD (Single Instruction Multiple Data) executes one operation across multiple data points packed in a register (for example, adding eight or sixteen integer values with a single instruction). Vector supercomputers pioneered this in the 1970s (such as the Cray-1), and the design reached commodity x86 and ARM processors through extensions like MMX, SSE, and NEON.

The trade-off is structural: every data lane must execute the exact same operation, and inputs benefit from contiguous, aligned layouts in memory.

Instruction
A single command the processor knows how to execute, like adding two numbers or loading one from memory.
Serial
Doing tasks one after another, each waiting for the previous.
Parallelism
Doing several operations at the same time.
02
The short version

A vector register holds a row of numbers side by side, one per lane. How many lanes you get is the width of the register divided by the size of your numbers.

Think of

the widened road

Add lanes rather than speed. Eight cars pass abreast under a single gantry that stamps all eight as they go, so the gantry still does one thing per pass and eight cars come out the other side with it done.

Widening the road

Think of data as traffic. A normal processor is a single‑lane road: one number at a time drives up to the point where the work happens, gets processed, and moves on. Raising the clock speed makes the cars go faster, which is the approach that ran into the heat ceiling. SIMD widens the road instead. Eight cars pass side by side under a single gantry that processes all eight at once.

Before a processor can touch a number, the number has to be in a register, a small and very fast storage slot on the chip itself. An ordinary register holds one number. A vector register is wider and holds a row of them side by side, and each slot in that row is a lane.

Lane count depends on how wide the register is and how big your numbers are, both measured in bits. An everyday whole number is commonly represented as a 32-bit value, so a 512‑bit register divides into sixteen lanes. Pixel values only run to 255, which fits in 8 bits, so the same register holds sixty‑four of them. That’s why image processing vectorises so well: the numbers are small, and a great many fit at once.

Figure 1: one register, divided

A vector register drawn as a single bar of fixed width, divided into equal lanes. Making the register wider or the numbers smaller divides the same bar into more lanes. 8 lanes256-bit register · AVX20123456732 bits a number · an everyday whole number

256bits

32bits

Instructions for one pass over the photo

4,500,000

64 lanes 562,500

Most modern phones support 128-bit vectors via ARM NEON. Desktop x86 CPUs standardise on 256-bit AVX2, while AVX-512 extends execution widths to 512 bits.

Modern GPUs take this execution model to an extreme scale with thousands of parallel lanes, typically structured as SIMT (Single Instruction, Multiple Threads).

Register
A tiny, fast storage slot inside the processor.
Vector register
A wide register that holds a row of numbers.
Lane
One slot in that row, holding one number.
Vector length
How many numbers a vector register holds at once.
03
The short version

Because one instruction targets the whole register, conditional logic requires masking off inactive lanes rather than skipping them in time.

Think of

the coned-off lane

The gantry still sweeps the full width on every pass. Cone off lanes three and seven and those two cars come out untouched, but the pass takes exactly as long as it did before.

When only some lanes should do work

Consider clamping values during brightness adjustment. In 8-bit unsigned arithmetic (0 to 255), adding a scalar directly can overflow: 250 + 20 wraps around to 14, turning bright pixels dark. Handling this conditionally requires selecting only pixels below 235.

In scalar code, a branch (if (val <= 235)) handles this per element. In SIMD, execution cannot branch on individual lanes within a register.

Instead, processors use masking (predication). A vector comparison instruction generates a bitmask or vector mask indicating which lanes meet the condition. Subsequent arithmetic instructions take this mask as an operand, only committing results to the active lanes. Clamping a byte is common enough that x86 and NEON both offer a saturating add as well, which stops at 255 without any mask at all; masking is the general tool, for the conditions the hardware has no dedicated instruction for.

Masking ensures correctness rather than saving cycles: the hardware still dispatches the instruction across the full vector width, discarding outputs from disabled lanes.

When lanes require genuinely divergent operations, execution serialises into multiple passes. If half the lanes require one operation and half require another, both instructions must execute with complementary masks. Frequent control flow divergence negates the throughput advantages of SIMD over scalar processing.

Figure 2: eight lanes, one of them coned off

Eight pixel values drawn as bars in one vector register, with a threshold line across them. Running the compare builds a mask, and the masked add moves only the lanes below the line. loaded · eight pixels, one register23521296255138244612331800 instructions issued255 − 20 = 235

20per channel

Lanes the add powers

8 of 8

6 lanes commit

Masking (predication)
Applying an instruction to only some lanes, using a row of on/off switches.
Divergence
When different lanes need to follow different logic, forcing multiple passes.
Zeroing / merging
What an inactive lane holds afterwards: zero, or its previous value.
04
The short version

Vector loads work best on neat rows of memory. Grabbing numbers scattered across different places takes extra work and slows things down.

When the numbers aren’t lined up

So far we’ve quietly assumed that the eight numbers you want next are sitting right next to each other. Loading a vector register in one move means grabbing a single unbroken block of memory, and that’s fast, because memory systems are built for exactly that.

Photos are lucky this way. They’re stored as one long run of pixel values, so a filter marching through them is always reading the next block along.

Not all data is arranged so neatly. If you have a list of user profiles where each entry holds a name, an email, and an account balance, the balances aren’t packed together; they’re separated by all the other fields. To load eight balances into a register, you can’t just scoop up an unbroken block.

Processors offer gather and scatter instructions for this. A gather instruction takes a list of separate memory addresses, fetches the value at each one, and packs them into a single vector register. Scatter does the opposite, writing the lanes back out to different locations.

The catch is speed. Grabbing numbers from eight different places still requires eight separate lookups under the hood. Hardware gather helps, but it is rarely as fast as reading a clean, uninterrupted block. In practice, keeping numbers side-by-side in memory often matters more than having wide registers to put them in.

Memory
The main store of working data, much larger and slower than registers.
Contiguous
Sitting immediately next to one another, in one unbroken block.
Gather
Filling a vector register from several separate locations at once.
Scatter
Writing the lanes of a vector register back out to several separate locations.
05
The short version

Real-world speedups usually land between 2× and 6×, because loading, storing, and housekeeping eat into the theoretical maximum.

The so-what

The theoretical ceiling is easy to state: eight lanes can be at most eight times faster, sixty‑four lanes at most sixty‑four. Real code lands well below that, because loading data, saving it back, and handling whatever leftover items don’t fit into an even chunk all take time. Two to six times faster is a normal result for work that suits SIMD.

Implementation

The brightness adjustment from the opening is small enough to write both ways. The input is a flat run of bytes in red, green, and blue order. The ordinary function visits one colour channel per loop. The vector version loads 32 channels at once, adds 20 to all of them with clamping, and writes them back. That’s ten full pixels plus two channels from the eleventh, but the boundaries don’t matter because every byte gets the exact same treatment.

naive.rsrust · one channel per iteration
pub fn brighten_naive(rgb: &mut [u8], amount: u8) {
    for channel in rgb {
        *channel = channel.saturating_add(amount);
    }
}
avx2.rsrust · 32 channels per vector
use std::arch::x86_64::*;

pub fn brighten_simd(rgb: &mut [u8], amount: u8) {
    if is_x86_feature_detected!("avx2") {
        unsafe { brighten_avx2(rgb, amount) }
    } else {
        brighten_naive(rgb, amount);
    }
}

#[target_feature(enable = "avx2")]
unsafe fn brighten_avx2(rgb: &mut [u8], amount: u8) {
    let step = _mm256_set1_epi8(amount as i8);
    let mut chunks = rgb.chunks_exact_mut(32);

    for chunk in &mut chunks {
        let ptr = chunk.as_mut_ptr().cast::<__m256i>();
        let values = unsafe {
            _mm256_loadu_si256(ptr.cast_const())
        };
        let brightened = _mm256_adds_epu8(values, step);
        unsafe { _mm256_storeu_si256(ptr, brightened) };
    }

    for channel in chunks.into_remainder() {
        *channel = channel.saturating_add(amount);
    }
}

The instruction names look intimidating, but each part just names a specific choice:

  • _mm256_ means we’re using a 256-bit register (wide enough for 32 bytes).
  • set1_epi8 copies our brightness step across all 32 lanes at once.
  • loadu and storeu grab 32 bytes from memory and put them back when done.
  • adds_epu8 is the saturating add from Section 3: it adds the brightness step to all 32 lanes in parallel and caps them at 255 so they never overflow.

The wrapper function checks whether the chip supports AVX2. If it does, it runs the fast vector loop in 32-byte chunks; if not, or for the few leftover bytes at the very end, it falls back to the plain loop.

07
The short version

Compilers can rewrite ordinary loops into vector instructions on their own, provided every iteration is completely independent.

The compiler was doing it all along

The result in the benchmark box is the most important takeaway: the plain, ordinary loop matched the hand-written SIMD code almost byte-for-byte. Nobody had to write special instructions for it. The compiler did it automatically.

This is called autovectorisation. When a compiler looks at a loop like for channel in rgb, it checks whether each step is doing the exact same independent math on consecutive items in memory. When that holds true, it quietly throws away the one-by-one loop and replaces it with 32-byte vector instructions.

The only catch is independence. If iteration two needs the answer from iteration one, or if items are scattered unpredictably in memory, the compiler plays it safe and leaves the loop alone.

You benefit from SIMD every day without ever writing it directly. Decoding images, streaming video, playing audio, unzipping files, and running local AI models all rely on it behind the scenes.

In practice, writing raw vector instructions by hand is rarely the first move. Write the clean, obvious loop, keep your data contiguous, and let the compiler widen the road for you.

Autovectorisation
The compiler rewriting an ordinary loop into vector instructions on its own.