This post summarizes the work I presented at the 21st International Conference on Availability, Reliability and Security (ARES 2026) in Linköping, Sweden. My name is József Sándor, and I’m a PhD student in the CrySyS Lab. PROPS is joint work with Roland Nagy and Prof. Levente Buttyán.
ROP Explained with Ransom Notes
Return-oriented programming (ROP) is a code-reuse attack that exploits memory corruption bugs. Real-world ROP techniques can be quite sophisticated, but the core principle is easy to grasp through the following analogy.
Imagine a crime film. The kidnappers want to send a threatening message, but they don’t want to write it by hand, because their handwriting could give them away. So they take a newspaper and cut out letters and words one by one, then glue them together into a ransom note.

ROP works in a surprisingly similar way:
- The newspaper is the memory space of the victim program.
- The letters and words are gadgets: short instruction sequences that already exist in the program’s code (the program itself and libraries like libc).
- Every letter and word in the newspaper is innocent, just as every gadget is legitimate code that the developers put there.
- The attacker selects the gadgets he needs and chains them together. The resulting payload carries a malicious intent, even though every single piece is innocent.
This is exactly why ROP bypasses XN (non-executable memory) protection: the attacker never injects new code, they only reuse what is already there.
On 32-bit ARM, a function typically returns via a pop {..., pc} instruction. After a buffer overflow, the attacker overwrites the saved return address and the area after it on the stack with the chain: the address of the first gadget, followed by the data and the addresses of the next gadgets. For example, to call system("/bin/sh"), the chain could pop the address of the string /bin/sh into r0, and then return into system().

Note that the chain lives on the top of the stack at the moment the vulnerable function returns. That is what we are going to look at.
Why should you care about ROP in 2026?
There is no ultimate defense against ROP, and this is especially true for legacy embedded systems. Modern ARM hardware features such as Pointer Authentication Code (PAC) and Branch Target Identification (BTI) are promising, but they are simply not available on devices built on ARMv7 and older architectures.
And those devices aren’t going away anytime soon. Imagine an embedded industrial environment with many devices, and a new memory corruption vulnerability comes to light. We know exactly where the bug is, but that doesn’t mean we can fix it. In these environments, hardware upgrades are impractical, vulnerabilities can’t always be patched promptly, and devices have to keep operating despite known memory-safety issues.
So the question we asked was: can we detect ROP attacks on such devices, without touching the existing software stack?
Design goals
We designed PROPS (Precision ROP Scanner) with the following goals in mind. It should:
- work on legacy systems with 32-bit ARMv7 and older architectures,
- be easy to deploy and compatible with existing systems, without modifying the existing software stack,
- detect ROP attacks rather than prevent them,
- use knowledge of a potential vulnerability (e.g., from a CVE database), so that we only need to watch the places that matter,
- learn only from benign stack content (anomaly detection), so that we don’t need to know in advance what future attacks will look like.
Overview of PROPS
PROPS has two phases.
Training phase (on a server). We take executable binaries, benign stack contents and memory mappings, extract features from them, and fit a PCA model.
Runtime inference (on the embedded device). The fitted PCA model is shipped to the device. When the monitored process reaches a pop {..., pc} we are interested in, we extract the features from its stack and ask the model whether it looks benign or like ROP.

The heavy lifting (training) happens off-device, and the device only needs to do cheap feature extraction and a PCA projection.
Turning the stack into features
The core question is: how do you describe a stack so that a ROP chain stands out, without knowing the exact chain?
Our answer is a simple heuristic:
- From the top of the stack, take 32 words of 4 bytes.
- Consider every 4-byte value as an address (they are 4-byte aligned on 32-bit ARM).
- Check where it points, look at the 40 bytes there (10 instructions), and give it a one-letter label:
| Label | Meaning |
|---|---|
F |
libc function |
S |
string |
N |
invalid instruction |
! |
SVC instruction (supervisor call) |
G |
pop {..., pc} instruction |
J |
branch-to-register instruction |
X |
executable region |
W |
writable region |
R |
readable region |
? |
address does not correspond to any mapped memory region |
So a stack becomes a string of 32 letters. A benign stack of /usr/bin/ls at a pop {r4, r5, r6, pc} looks like a rather random mixture of W, ?, X, and so on. It is mostly data, saved registers, pointers to writable memory and values that don’t point anywhere.
A ROP chain, on the other hand, is made of addresses that point to gadgets, functions and strings. For example, the chain for system("/bin/sh") is encoded as G S ? F: a gadget, the address of the string, a filler value, and the address of a libc function. When this chain is written on the top of the stack, it overwrites part of the benign stack. Its position in the 32-letter window (the offset) depends on the vulnerable function.

Letters like F, S and G tend to appear in a very specific combination in an exploit and rarely show up like that in normal stacks. That is the signal that we want to catch.
ROP as an anomaly
We treat detection as an anomaly detection problem, in a semi-supervised setting: we only learn what benign stacks look like.
- We fit a Principal Component Analysis (PCA) model on the benign feature vectors.
- At runtime, we project the stack’s feature vector to the low-dimensional space, reconstruct it back, and measure the reconstruction error.
The intuition: PCA learns the directions along which benign stacks vary. A benign stack can be reconstructed well from these directions, so its error is low. A stack that contains a ROP chain has a pattern that the model has never seen, so it is reconstructed poorly, and the error is high.
The figure below shows this on a toy 2D example.

If a stack’s reconstruction error goes above a predefined threshold, we flag it as ROP.
Data: benign stacks and ROP chains
We collected the benign data on a Raspberry Pi 3 Model B running the 32-bit version of Raspberry Pi OS Lite:
- We ran 900+ programs under GDB.
- We set a breakpoint at
main(), where usually every library is already loaded, and made a full memory dump. - We set breakpoints at every
pop {..., pc}instruction, and saved the top of the stack (32 × 4 bytes) whenever one was hit.
For the malicious side, we generated ROP chains with angrop (Zeng et al., NDSS 2026):
- 60 selected common ELF binaries (
ls,grep,find, …), - 3 attacker objectives: spawning a shell, creating a new executable page, making the stack executable,
- 4 methods: calling functions or system calls, each with gadgets either from libc or from the target program,
- 2 modes: ARM and Thumb.
Evaluation
We trained on 80% of the benign data. We used ROP chains only for testing. The results:
- a true positive rate of ~99%,
- a false positive rate of about 2-3%.
In other words, the model catches nearly all ROP chains, while raising a false alarm on only a small fraction of benign stacks.
Proof-of-concept implementation
Detection is only useful if it can run on the device, so we built a proof of concept for Linux:
- We place trigger points (implemented via uprobes) on some return instructions of the monitored process (typically in the “vulnerable function” that we learned about, e.g., from a CVE).
- A loadable kernel module extracts the necessary data from the running process. When the execution reaches a trigger point, the kernel module gets the control, extracts the top of the stack and the memory mapping information, and sends them to the inference system.
- The monitored process runs further.

Because there is no change to the monitored program (or to the rest of the software stack), this fits our “easy to deploy” goal.
Runtime overhead
The monitored process is briefly paused while the kernel module collects the data. The processing time of the module is fixed, except for the traversal of the memory mappings, so the per-invocation latency is low.
We measured the runtime overhead on 60 common utility programs with 1, 3, 6 and 12 trigger points placed at different function return points. Keep in mind that the same trigger point can be reached multiple times. With 1-3 trigger points, the average overhead is about +2%, however, it increases with the number of trigger points.
This is why PROPS monitors only the return points of known (or potentially) vulnerable functions, instead of every return in the system.
Validation on real vulnerabilities
To check that this also works outside of our artificially generated dataset, we tested PROPS on two real vulnerable programs, using ROP chains from Exploit-DB:
- Crashmail 1.6 (CVE-2018-25223)
- PMS 0.42 (CVE-2018-25224)
In both cases, PROPS successfully detected the exploitation.
Conclusion and what’s next
To sum up, PROPS is a ROP detection method for legacy ARM-based devices. It is easy to deploy and compatible with existing systems, it monitors only the return points of known (or potentially) vulnerable functions, it reaches a high detection rate (99% TPR), and the runtime overhead is practical with 1-3 trigger points.
We are currently working on:
- validation and runtime measurements on additional vulnerable programs,
- reducing the false positive rate via a ROP chain integrity check.
For further details, you can check our paper here.