Anti-Debugging for Noobs, Part 1: How a Program Notices the Debugger
First principles of anti-debugging on Linux: the one-tracer rule, ptrace self-trace, TracerPid in /proc, parent checks, timing gaps and 0xCC breakpoint scanning, plus how analysts answer each check.
Anti-Debugging for Noobs, Part 1: How a Program Notices the Debugger#
Anti-debugging is the set of tricks a program uses to answer one question: "is someone watching me run?" Packers, DRM systems, banking trojans and crackmes all ask it. Reverse engineers ask the opposite: "how do I watch without being noticed?"
This first part covers the five checks you will meet in almost every Linux sample, why each one works, and how an analyst answers it. Part 2 goes deeper into self-integrity checks and anti-VM tricks.
The one rule everything builds on#
Linux debugging runs through ptrace. The kernel allows exactly one tracer per process. If a debugger is already attached, a second ptrace attach fails. Almost every beginner anti-debug check is a creative way of asking the kernel: "is there already a tracer on me?"
analyst runs: gdb ./crackme | v +-------------------------------+ | process starts under a tracer | +-------------------------------+ | v check 1: does PTRACE_TRACEME fail? -> exit or misbehave check 2: is TracerPid non-zero? -> exit or misbehave check 3: is my parent gdb/strace? -> exit or misbehave check 4: did this step take ages? -> exit or misbehave check 5: is there 0xCC in my .text? -> exit or misbehave | v the real payload runs only if every check passes
Keep the diagram in mind; the rest of the post walks down each line.
Check 1: tracing yourself first#
A process can ask the kernel to trace itself with PTRACE_TRACEME. If a debugger already traces it, the request fails with -1. That single return value is the oldest anti-debug check in the book:
#include <stdio.h> #include <sys/ptrace.h> int main(void) { if (ptrace(PTRACE_TRACEME, 0, 1, 0) == -1) { puts("debugger spotted"); return 1; } puts("running clean"); return 0; }
A side effect helps the attacker twice: once a process traces itself, nobody else can attach later. gdb -p $(pidof crackme) comes back with "Operation not permitted".
The analyst's answers, from light to heavy:
- Start the binary inside gdb instead of attaching later;
PTRACE_TRACEMEfrom inside a traced run succeeds because the tracer slot is the same one. - Patch the check out once you find it in the disassembly; it is usually a single conditional jump.
LD_PRELOADa shim that wrapsptraceand returns success forPTRACE_TRACEMEwhile doing nothing.
Check 2: reading your own report card#
The kernel publishes the tracer's PID in /proc/self/status as the TracerPid field. Zero means free, anything else means watched:
#include <stdio.h> #include <string.h> static int being_traced(void) { FILE *f = fopen("/proc/self/status", "r"); char line[256]; while (fgets(line, sizeof(line), f)) { if (strncmp(line, "TracerPid:", 10) == 0) { int pid = 0; sscanf(line + 10, "%d", &pid); fclose(f); return pid != 0; } } fclose(f); return 0; } int main(void) { if (being_traced()) { puts("nice try"); return 1; } puts("running clean"); return 0; }
Try it:
./tracerpid_check # prints "running clean" strace ./tracerpid_check # prints "nice try", strace is a tracer too
This check catches strace, ltrace (which also uses ptrace), and gdb alike. The analyst's answers:
LD_PRELOADa shim aroundfopen/fgetsthat rewritesTracerPid:lines to zero.- In gdb, break after the read, flip the comparison result, continue.
- Patch the branch; again it is one conditional jump in the disassembly.
Check 3: who is your parent#
A binary launched from a shell has a shell as its parent. A binary launched from gdb has gdb as its parent. Reading /proc/self/stat gives the parent PID, and /proc/<ppid>/comm gives its name:
cat /proc/self/stat | awk '{print $4}' # ppid field cat /proc/$(ps -o ppid= -p $$)/comm
The check compares the parent's name against a denylist: gdb, strace, ltrace, r2, radare2. More careful samples hash the name instead of storing the strings in cleartext, because strings output with "gdb" in it is a gift to the analyst.
The analyst's answers are boring and effective: rename the debugger binary, or run the target under a launcher that double-forks so the immediate parent is init or a clean wrapper.
Check 4: time is the loudest witness#
Single-stepping through code is millions of times slower than running it. A program that measures its own execution time sees the debugger as a giant gap between two timestamps:
#include <stdio.h> #include <time.h> static long now_ns(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1000000000L + ts.tv_nsec; } int main(void) { long a = now_ns(); volatile long x = 0; for (long i = 0; i < 1000; i++) x += i; long b = now_ns(); if (b - a > 1000000L) { /* a plain loop should take far less */ puts("too slow, someone is stepping"); return 1; } puts("running clean"); return 0; }
On x86, samples often use the rdtsc instruction instead, which counts CPU cycles and cannot be faked by lying about the wall clock. Either way, the analyst's answer is to keep the run fast at the measured spot: break before the first timestamp and after the second, or patch the threshold, or use an emulator that can pause time itself.
Check 5: hunting your own breakpoints#
A software breakpoint is not a flag in some register. gdb writes the byte 0xCC (int3) over the first byte of the instruction where you set the breakpoint, and restores the original byte when the trap fires. That means a program can find its own breakpoints by reading its own code:
#include <stdio.h> extern char _etext; /* linker symbol: end of .text */ extern char _stext; /* start of .text */ static int count_int3(void) { int n = 0; for (char *p = _stext; p < _etext; p++) if ((unsigned char)*p == 0xCC) n++; return n; } int main(void) { if (count_int3() > 0) { puts("found breakpoints in my own code"); return 1; } puts("running clean"); return 0; }
Real samples do not scan everything; they checksum a small function and compare against a compile-time constant. Same idea, less noise.
The analyst's answers:
- Use hardware breakpoints (
hbreakin gdb). They live in debug registers DR0-DR7 and leave the code bytes untouched, so the scan finds nothing. - Patch the scan, once you spot the tight loop comparing bytes against
0xCC.
Signal games#
When a traced process hits int3, the kernel delivers SIGTRAP. A program can install its own SIGTRAP handler and deliberately execute int3; if the handler runs, all is well, but if a debugger swallows the trap first, the handler never runs and the program knows. A softer version just checks whether a signal it raised got reported back through waitpid semantics by a cooperating child. The details vary, the principle does not: unexpected silence on a channel the program itself controls is a detection.
Analysts answer by configuring gdb to pass the signal through: handle SIGTRAP nostop noprint pass.
A note on debug registers#
Hardware breakpoints use the DR registers, and a process cannot read them directly from userspace, but ptrace(PTRACE_PEEKUSER) against the debug register area lets a traced process ask the kernel for its own DR values. Non-zero DR0-DR3 means hardware breakpoints exist. Some packers clear them via PTRACE_POKEUSER before continuing, which is both a check and a sabotage in one syscall.
The analyst's general playbook#
Every check above falls to one of four moves:
- Avoid the check. Start inside the debugger instead of attaching; use hardware breakpoints instead of software ones.
- Lie to the check.
LD_PRELOADshims that fakeptrace,fopen, or timing results. - Patch the check. Flip one conditional jump; the binary stops asking.
- Change the ground truth. Run under an emulator (QEMU, unicorn) where the analyst controls time, registers, and memory below the program's feet.
Pick the cheapest move per check. Patching is permanent but requires finding the check first; shims are quick but only affect dynamic calls.
Reading a check in the disassembly#
Knowing the five checks is one thing; spotting them in a stripped binary is the real skill. The signatures to scan for:
# syscall numbers and magic constants give checks away objdump -d ./crackme | grep -i "0xcc" objdump -d ./crackme | grep -i ptrace strings ./crackme | grep -i "tracerpid\|/proc/self\|gdb\|strace"
A ptrace call in a program that has no business debugging anything is a finding. A tight loop comparing bytes against 0xCC is a finding. A call to fopen with /proc/self/status as the argument is a finding. Mark each site in your notes before patching anything, because packers often place the same check in three spots and fire the payload only when all three report clean.
Combining checks: the honest cost model#
One check is a speed bump. Three checks in different styles are a real budget:
- A static
LD_PRELOADshim answers the library-level checks but does nothing against inline syscalls. - Patching answers everything but costs analysis time per site.
- An emulator answers everything at once but costs setup time and breaks on programs that probe for the emulator itself.
Malware authors know this math, which is why real samples mix categories: one ptrace check, one timing check, one checksum. Expect the mix, not the single trick.
2026 note: the checks still work, the surroundings hardened#
The five checks above remain valid on a 2026 kernel; what changed is that the kernel and the toolchain hardened the defaults around them:
user | v [ shell / systemd ] <-- ptrace_scope, Yama, LSM policies | v [ anti-debug binary ] ----- ptrace /proc/self/status rdtsc 0xCC scan | | | v | kernel: at most one tracer | | | v +--- traced run: TracerPid non-zero, timing deltas visible
- Yama
ptrace_scope(/proc/sys/kernel/yama/ptrace_scope) defaults to 1 on mainstream distros: a process may only trace its own children or processes that explicitly opted in. That removes casual attach-based loggers and casual anti-anti-debug without preparation; starting the process inside the debugger matters even more. - Seccomp and user-namespace boundaries filter the syscall neighbors anti-debug samples rely on -
ptrace,process_vm_readv,perf_event_open. Inline-syscall checks bypass theLD_PRELOADshims that wrap those calls, so cost per check goes up. - eBPF observability is the new kernel-side surface an analyst owns. The target closes the ptrace slot by binding it to itself, but an eBPF observer on
tracepoint:syscallsstill sees the same calls from outside the target's view. If the target fights back with eBPF programs of its own,perf_event_paranoidand eBPF LSM are the referee. - Timing checks stay noisy;
rdtscstill counts cycles inside a VM, but the hypervisor binds time to the virtual address space, and an emulator already owns the clock. A timing check tells you when a three-layer setup is warranted; it is not the last word on its own.
Practical takeaway: in the first twenty minutes of a 2026 binary analysis, put eBPF tracepoints, the ptrace signature, and TracerPid reads side by side. If all three point at the same site, the packer placed the same check three times, and patching one will trip the other two.
Think of it as a trust boundary#
It helps to name the theme the five checks share: each one asks the kernel or a side channel to keep a consistency promise on the user's behalf. The one-tracer rule is the kernel's promise; TracerPid is that promise made visible; the parent check is a userspace guess about a promise the kernel never made; the timing bound leans on layers above the kernel - the hypervisor and the emulator; and the 0xCC scan is the program asking a debugger-modified image to report on itself. So write down which layer each check stands on. If a check leans both on the kernel's word and on the debugger's good faith, the attacker usually walks in through the weakest layer, and hardening one side leaves the other untouched.
Where to draw the line#
The line between a check and a patch is what the check leans on. If it leans on the kernel's word and you cannot relax the kernel setting, a debugger configuration or a small patch is usually the cheapest move; if it leans on debugger behaviour, renaming the debugger binary or switching to hardware breakpoints often suffices. In 2026, eBPF observability sits between the two ends: independent of the target yet anchored in the kernel, so it caps both techniques. Do not start patching before you write down which layer each check holds onto; pick the cheap move from the right layer.
| Check | Layer it leans on | Cheapest analyst move |
|---|---|---|
PTRACE_TRACEME | One-tracer rule (kernel) | Start inside the debugger or use a shim |
TracerPid read | /proc publication | Shim the read, then patch |
| Parent name | Userspace guess | Rename the binary, launch via double-fork wrapper |
| Timing gap | Hypervisor or emulator setup | hbreak between the timestamps, patch the threshold |
0xCC scan | Debugger behaviour | Hardware breakpoints (hbreak) |
SIGTRAP silence | Kernel signal delivery | handle SIGTRAP nostop noprint pass |
Read the table backwards as a packer: every row raises the cost of whichever single technique that row targets. Mixing layers keeps the analyst from dropping the whole net with one shim or one patch. Running them together in one binary is exactly what makes them reinforce each other. Real samples run most of these rows inside the same binary, each with its own trigger threshold and its own reporting channel.
Lab exercises#
- Compile the
PTRACE_TRACEMEsample. Run it plain, under gdb, and under strace. Note the differences. - Write the
TracerPidcheck and watch/proc/self/statusby hand while it runs under strace. - In gdb, set a breakpoint inside the timing-check binary, single-step a few instructions, and observe the check fire. Then break before and after instead, and watch it pass.
- Build the
0xCCscanner, set a software breakpoint inmainunder gdb, run, and count the found bytes. Repeat withhbreak. - Write the
LD_PRELOADshim forptraceand defeat check 1 without touching the binary.
Where this leaves you#
You can now name the five checks, explain the one-tracer rule that powers the first three, and pick an analyst move per check. That vocabulary is the entry ticket for every anti-debug writeup and crackme walkthrough you will read from here on.
Lab upkeep#
One thing to internalise before every crackme run: packers spread the same check across several sites and gate the payload on all of them reporting clean. A single bypassed check is not the cause when the payload refuses to run; treat each check site as a separate, marked patch.
Layered checks: self-checksum, gated decryption, timing, anti-VM#
Mechanical checks catch notice; layered checks change the outcome when one is bypassed. The four that recur in post-2020 samples:
-
Self-checksum over
.text. The binary FNV-hashes its own code and compares against a baked value. A patched byte flips the hash. The analyst counter is to keep the running copy consistent, or pre-compute the expected hash and patch the comparison target. If the hash is compared to an immediate, patch the jump-equality withjz/jnzflip, not the check itself. -
Gated decryption. A section stays as ciphertext until every gate passes: no debugger, expected timing delta, matching checksum, sane DMI. The payload derives a key from those checks and
mprotect+decrypts only if all hold. This converts "did you notice" into "did you fix every gate". The reliable lab counter is to run the binary under a hypervisor or a VM with CPU patches, or mark each gate return with a Frida hook (Interceptor.attachon the check function) and return0. -
Per-CPU timing loops. Two or more
rdtscpairs wrap every action; any preempt is treated as a debugger. Samples also trysched_getcpuplus expected latency tables. Counter: virtualise the TSC in your analysis VM or uselibdisasm/ptracestop points around the loops so the loop sees no delta. -
Anti-VM probes. DMI strings in
/sys/class/dmi/id/, CPUID hypervisor bit, MAC OUI, QEMU/VBOX services. Counter: patch sysfs/DMI with a modifieddmimodule, use KVM host-only mode with hypervisor bits clear, or select a custom CPUID via the VM profile.
Treat every check as a pivot. A gdb one-liner that marks a successful gate as "clean" is usually enough for one branch; the layered binary will re-check. The lab write-up should record each patch site and each frame that tripped it.
What part 2 covers#
Part 2 moves past single checks: self-checksums over whole sections, code that decrypts itself only if no check fired, timing loops tuned per CPU, anti-VM probes through /sys and CPUID, and the gdb scripts that cut through them one by one. If the five checks here feel mechanical, the layered ones next time will feel like chess. Read the checks as a set inside one binary, not as isolated tricks.
What do you think?
React to show your appreciation