Let’s add two numbers
Here’s a complete C function that adds two unsigned 128-bit integers:
typedef unsigned _BitInt(128) u128;
u128 add(u128 a, u128 b) {
return a + b;
}
Now let’s compile it for 32-bit RISC-V:
clang --target=riscv32 -march=rv32imac -mabi=ilp32 -O2 -S add.c -o -
You can see the output on Compiler Explorer, alongside versions compiled with GCC and with the Zicond extension we’ll look at below.
Here’s the relevant snippet:
add a1, t1, t0
add a2, a7, t2
sltu t0, a1, t1
add a2, a2, t0
beq a2, a7, .LBB0_2
sltu t0, a2, a7
.LBB0_2:
add a5, a5, a3
add a4, a4, a6
add t0, t0, a5
sltu a3, a5, a3
sltu a5, t0, a5
Wait, why is there a beq in an addition? That’s a conditional branch, right?
RV32 registers hold 32 bits, so LLVM has to split our 128-bit addition into four smaller ones, passing the carry from each to the next.
This is how the compiled code performs the addition; a0 and b0 are the lowest 32-bit words of our inputs; a1 and b1 are the next ones.
All values are unsigned, low32() keeps only the lowest 32 bits, and comparisons return 0 or 1.
sum0 = low32(a0 + b0)
carry0 = (sum0 < a0)
sum1 = low32(a1 + b1 + carry0)
if sum1 == b1:
carry1 = carry0
else:
carry1 = (sum1 < b1)
Can you guess why sum1 == b1 is checked?
CPUs such as x86 and AArch64 have instructions to perform conditional moves (cmov), allowing carry propagation to be implemented without a branch.
But on RV32, even comparing two 64-bit integers with < produces a branch.
What about the usual bit mask?
We’ve been talking about carry propagation in large integers, but if you’ve written constant-time code, you’ve probably used some version of this everywhere:
#include <stdint.h>
uint32_t ct_select(uint32_t bit, uint32_t a, uint32_t b) {
uint32_t mask = -(bit & 1);
return (a & mask) | (b & ~mask);
}
If the low bit of bit is set, mask is all ones, so the expression keeps a and zeros out b.
Otherwise, the mask is zero and we get b.
Just bitwise operations, no branch in the source.
Let’s compile that for RV32 with clang 23:
ct_select:
andi a0, a0, 1
beqz a0, .LBB0_2
mv a2, a1
.LBB0_2:
mv a0, a2
ret
Aaaahhhhhhhh, a beqz instruction, branching on the bit we just masked. So much for carefully writing the selection with bitwise operations.
And this also happens on 64-bit RISC-V.
You can see the compiled code on Compiler Explorer which includes both targets, plus clang 17, GCC and Zicond for comparison.
Why include clang 17? Because the branch was there in 15, gone in 16 and 17, and back from 18 to 23. Fun, uh?
So, even if you reviewed assembly code with a given version of the compiler, and everything looked fine, every change to the compiler version of compiler flags requires a new review.
What about Zig?
Let’s try the 128-bit addition in Zig, along with the same bit-mask selection:
export fn add(a: *const u128, b: *const u128, r: *u128) void {
r.* = a.* +% b.*;
}
export fn ctSelect(bit: u32, a: u32, b: u32) u32 {
const mask = 0 -% (bit & 1);
return (a & mask) | (b & ~mask);
}
zig build-obj add.zig -target riscv32-freestanding -O ReleaseFast -femit-asm=add.s -fno-emit-bin
Same beq after the second word, and the selection gets the same beqz (Compiler Explorer code).
Changing the source language doesn’t get us out of this. And yes, Rust has the same issue.
What about other platforms?
Now let’s compile C examples for a few other targets.
I also added a 64-bit a < b comparison, since that was enough to produce a branch on RV32.
These are the numbers of conditional branches and conditional returns emitted by clang 23 at -O2.
As usual everything can be verified on Compiler Explorer:
| Target | add |
ct_select |
64-bit < |
|---|---|---|---|
| x86_64, x86 | 0 | 0 | 0 |
| AArch64, 32-bit ARM, Cortex-M3 | 0 | 0 | 0 |
| MIPS32, LoongArch64, WebAssembly | 0 | 0 | 0 |
| Cortex-M0 | 0 | 1 | 1 |
| 32-bit PowerPC | 0 | 1 | 2 |
| RV32 | 1 | 1 | 1 |
| RV64 | 0 | 1 | 0 |
| RV32 and RV64 with Zicond | 0 | 0 | 0 |
Targets with zeros are safe. Everything else has ugly side channels in spite of source code looking like it runs in constant time.
WebAssembly has a select (cmov) instruction, so no obvious conditional jumps are visible in the modules, but then WebAssembly compilers can do whatever they want. On platforms without equivalent native instructions, it’s likely that we’ll get a jump.
Cortex-M0 (Thumb-1) and generic 32-bit PowerPC don’t have an cmov-like instructions, so they branch.
Let’s try GCC
Now here’s a pleasant surprise: GCC 16.1 compiles both examples without branches on RISC-V. Its carries use sltu, and it leaves the mask arithmetic alone.
Cool. But let’s make a small change: derive the mask from a comparison.
uint32_t m = -(uint32_t) (x < y);
return (a & m) | (b & ~m);
And… the branch is back!
GCC now emits a bgeu on both RV32 and RV64 (Compiler Explorer).
It also branches on 64-bit comparisons on RV32.
Can we hide the mask from the optimizer?
For the bit-mask example, there’s a common workaround: pass the mask through an empty asm statement before using it. Let’s do that:
uint32_t ct_select(uint32_t bit, uint32_t a, uint32_t b) {
uint32_t mask = -(bit & 1);
__asm__("" : "+r"(mask));
return (a & mask) | (b & ~mask);
}
The assembly does nothing, but its declaration tells the compiler that it may change mask.
Now, both versions compile without branches on RV32 and RV64. Phew.
Can we do the same for the addition?
I tried hiding the inputs behind a memory barrier, and the branch stayed.
But putting a register barrier on each of their 32-bit words worked, and every carry became an sltu.
Both attempts are on Compiler Explorer.
To be honest, I wouldn’t rely on either barrier experiment as a fix for the addition.
For arithmetic involving secrets, I’d avoid integer types wider than two registers and write the carries explicitly, using 32-bit words on RV32.
Currently, clang 23 keeps my hand-written carry chain free of branches, but who knows what will happen in the next releases.
Giving LLVM the missing instructions
For RISC-V, there’s a solution, though: RISC-V has an extension called Zicond.
It adds czero.eqz and czero.nez, which zero a register depending on whether another register is zero.
And that can be used to select a value without branching.
Let’s enable it with -march=rv32imac_zicond and compile our bit-mask example again:
ct_select:
andi a0, a0, 1
czero.eqz a1, a1, a0
czero.nez a0, a2, a0
or a0, a0, a1
ret
Yay, no jumps. Every case tested above is free of branches with Zicond enabled.
Zicond is part of the RVA23 profile, but unfortunately many cores in use today don’t implement it, especially microcontrollers.
And even if it’s available, there’s an important detail that’s easy to overlook: the Zicond specification only guarantees that their timing is independent of the data if the Zkt extension is also implemented.
Writing secure, portable code is hard. Protecting against side-channels is as footgunish as zeroing secrets.
Oh, and if you haven’t read it yet, Thomas Pornin’s Why constant-time crypto? and Constant-time multiplication pages are absolutely worth a read.