Research/memfd-fileless-execution-research
MalwareCVE-2019-5736InfoPublic

The binary that wasn’t there

Linux will run a program that has no name. I went looking for folklore and found memfd_create: anonymous RAM, an ELF, and an exe symlink that has been willing to testify the entire time. Stock Debian still leaves the door open.

Name
The binary that wasn’t there
Type
Malware reverse-engineering
CVE
CVE-2019-5736
CVE Risk
informational
Disclosure Status
public
Vendor
Linux kernel / Debian
Affected
Linux 3.17+; stock Debian still ships vm.memfd_noexec=0. Container hosts must allow-list runc.
Published
20 Aug 2026
Updated
01 Sept 2026
Tags
linux, memfd, fileless, detection, debian, hardening, notes

Linux will run a program that has no name

I wrote that sentence on a sticky note and waited for it to feel like folklore. It never did. It is a consequence of how the kernel thinks about files, and once you see the object model the magic evaporates in the most satisfying way.

The claim, stated carefully, because I have already had this argument in Slack:

  • memfd_create() allocates an anonymous file whose pages live in RAM.
  • An ELF image can be written into that file descriptor.
  • The kernel will execute it via /proc/self/fd/N, fexecve(), or execveat(..., AT_EMPTY_PATH).
  • There is no directory entry on a persistent filesystem. The ELF loader does not care.

The riddle is not “how does Linux run code without a file?” Linux never required a path. It required a struct file. Everything else is commentary.

This is the teaching pass. Mechanism, a benign lab check using /bin/echo, Debian-side defense, detection, and what it means when the thing in RAM is a reverse shell or a C2. I am not publishing an implant. I am publishing the map.

What memfd_create actually allocates

memfd_create(2) arrived in Linux 3.17. It returns a file descriptor to an anonymous file. The file behaves like a regular file: read, write, ftruncate, mmap, seals. Its backing store is anonymous memory, with the same lifetime rules as mmap(..., MAP_ANONYMOUS). When the last reference drops, the pages go with it.

The name argument is a joke the kernel tells /proc. Diagnostic only. In /proc/<pid>/fd/ you will see something like /memfd:ghost (deleted). That is not a path you can open() from a normal mount. There is no dentry on ext4 or xfs. The inode is real. The phone book entry is not.

Useful flags, because the ABI grew up in public:

Flag Role
MFD_CLOEXEC Close-on-exec. Sensible for a one-shot loader.
MFD_ALLOW_SEALING Enables fcntl seals (F_SEAL_WRITE, F_SEAL_SHRINK, and friends).
MFD_EXEC Explicitly create an executable memfd.
MFD_NOEXEC_SEAL Create non-executable and seal the execute bit so it cannot be added later.
MFD_HUGETLB Huge-page backing. Not the point of this note.

The descriptor is O_RDWR. Historically the underlying file also came with the execute bit set. That combination - writable and executable - is the sort of thing kernel people have been trying to make socially unacceptable for a decade. ChromeOS and others had the bruises to prove it. Hence the later flags, and hence vm.memfd_noexec.

Why the ELF loader does not ask for a name

binfmt_elf operates on a struct file, not on a pathname. A path is only the user-space ritual for finding the file. Once the kernel has the file object, it maps segments, fixes up the interpreter, and transfers control.

Three equivalent ways to hand over the file:

  1. execve("/proc/self/fd/N", argv, envp) - procfs resolves the symlink to the same file.
  2. fexecve(fd, argv, envp) - glibc’s wrapper.
  3. execveat(fd, "", argv, envp, AT_EMPTY_PATH) - no /proc required, no pathname required.

The last form is the cleanest statement of the punchline. You do not even give it a name with extra steps. You give it a descriptor and an empty string.

After a successful exec:

  • /proc/<pid>/exe commonly reads /memfd:<name> (deleted).
  • /proc/<pid>/maps shows executable mappings backed by anon_inode:[memfd] (or equivalent).
  • argv[0] is whatever the caller supplied. It need not match anything on disk. It can pretend to be nginx. The exe symlink will not play along.

Shebang scripts are the fragile case. MFD_CLOEXEC can close the interpreter’s view of the script. ELF is the reliable case, which is why every serious discussion of this technique ends up talking about ELF.

noexec on a normal mount does not apply. There is no such mount. Permission to map executable pages is a property of the memfd and of the process, not of /tmp’s mount flags. That is the entire reason this technique is interesting when /tmp has been spanked with noexec. I wrote that last sentence twice. It is the whole paper.

Lab demonstration - prove the kernel will do it

I wanted a demo that does not invent a payload. So I copied /bin/echo - a binary the system already trusts - into a memfd and replaced the process with that image. The bytes that run never exist as a named file created by this process. That is the whole trick, isolated from everything interesting an implant might do afterward.

Run this only on a host you own or are authorized to use. It is a kernel-behavior check, not a remote-access sample. A reverse shell is a different program. This note does not include one.

C
#define _GNU_SOURCE
#include <errno.h>
#include <fcntl.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mman.h>
#include <unistd.h>

static int copy_file_to_fd(int src, int dst)
{
    char buf[4096];
    ssize_t n;

    while ((n = read(src, buf, sizeof(buf))) > 0) {
        char *p = buf;
        ssize_t left = n;
        while (left > 0) {
            ssize_t w = write(dst, p, (size_t)left);
            if (w < 0) {
                if (errno == EINTR)
                    continue;
                return -1;
            }
            p += w;
            left -= w;
        }
    }
    return (n < 0) ? -1 : 0;
}

int main(void)
{
    int src = open("/bin/echo", O_RDONLY);
    if (src < 0) {
        perror("open /bin/echo");
        return 1;
    }

    int fd = memfd_create("demo", MFD_CLOEXEC | MFD_EXEC);
    if (fd < 0) {
        /* Older kernels: omit MFD_EXEC */
        fd = memfd_create("demo", MFD_CLOEXEC);
    }
    if (fd < 0) {
        perror("memfd_create");
        close(src);
        return 1;
    }

    if (copy_file_to_fd(src, fd) != 0) {
        perror("copy");
        close(src);
        close(fd);
        return 1;
    }
    close(src);

    if (lseek(fd, 0, SEEK_SET) < 0) {
        perror("lseek");
        close(fd);
        return 1;
    }

    char *argv[] = { "echo", "executed-from-memfd", NULL };
    char *envp[] = { NULL };

    if (fexecve(fd, argv, envp) < 0) {
        perror("fexecve");
        close(fd);
        return 1;
    }

    /* not reached */
    return 0;
}

Build and run:

Shell
gcc -O2 -o memfd_exec_demo memfd_exec_demo.c
./memfd_exec_demo

A successful run prints executed-from-memfd. That string came from echo, whose image lived only in the anonymous file at the moment of exec. I sat back. The kernel did exactly what the man page said it would do. The folklore was a man page.

If you need the form that does not require /proc or glibc’s fexecve:

C
#include <sys/syscall.h>
#include <linux/fcntl.h>  /* AT_EMPTY_PATH */

if (syscall(SYS_execveat, fd, "", argv, envp, AT_EMPTY_PATH) < 0)
    perror("execveat");

What a second terminal should see if you catch the process mid-flight (easier if you replace echo with something that sleeps):

  • /proc/<pid>/exe -> /memfd:demo (deleted)
  • /proc/<pid>/maps with executable mappings on anon_inode:[memfd]
  • argv[0] equal to echo, which is a lie the process told about itself

If vm.memfd_noexec=2 is set in the namespace and the create flags are not an allowed combination, memfd_create fails. That is the hardening knob doing its job, not a broken demo.

The loader in this listing still reads /bin/echo from disk. That is intentional: it proves the executed image is the memfd, while refusing to smuggle a custom ELF into the brief. For detection practice, the interesting artifact is the exe path after fexecve, not where the bytes were copied from.

The kernel grows a conscience - vm.memfd_noexec

For years, executable was the default. Convenient for runc. Inconvenient for everyone trying to keep W^X meaningful.

Linux 6.3-era work added:

  • MFD_EXEC and MFD_NOEXEC_SEAL so the caller states intent at creation time.
  • A PID-namespace sysctl, vm.memfd_noexec, with three values:
Value Meaning
0 Implicit MFD_EXEC. Historical default. What you still get on a stock Debian kernel unless you change it.
1 Implicit MFD_NOEXEC_SEAL. Callers that need an executable memfd must pass MFD_EXEC.
2 Reject creates that do not pass MFD_NOEXEC_SEAL. High-assurance. Easy to break things.

The sysctl is hierarchical. The most restrictive ancestor namespace wins. A child cannot loosen the parent. That is the correct shape for a security knob, and it is also why you test before setting 2 on a container host. I will die on this hill. It is a small hill. It has a good view of runc.

runc is the legitimate user that forced the design to keep an executable path at all. After CVE-2019-5736, runc copies itself into a sealed memfd so a compromised container cannot overwrite the host /usr/bin/runc through /proc/self/exe. Current runc tries MFD_EXEC first. Detection engineers who alert on "any memfd exec" will page themselves every time a pod starts, which is a lifestyle, not a strategy.

Debian, as it actually ships

Assume a modern Debian (13 Trixie or later): systemd, AppArmor on by default, /tmp as tmpfs, auditd not installed unless someone asked for it, vm.memfd_noexec=0.

That last clause is doing a lot of work. The distribution gives you an LSM, a service manager with excellent sandboxing directives, and a kernel that will still cheerfully create executable memfds unless you opt into the new policy. If that is uncommon, I would like to see their common.

What the stock baseline does not stop

AppArmor is path-centric. An anonymous inode is not a path. A well-written profile can still stop a confined daemon from exec’ing anything new, which is the actually useful property. An unconfined user or an unconfined daemon is unimpressed.

/tmp as tmpfs is a hygiene change, not a memfd control. /dev/shm is usually nosuid,nodev and too often still executable. If you only block memfd, the next conversation is O_TMPFILE on an executable tmpfs. I have that tattooed on the inside of my eyelids now.

The one host-wide knob worth turning first

Plain text
# /etc/sysctl.d/30-memfd-noexec.conf
vm.memfd_noexec = 1

Then sysctl --system. Soak it. Anything that needed an executable memfd and forgot to pass MFD_EXEC will complain. runc already knows the dance. 2 is for hosts that do not run container engines, desktops, or mystery vendor agents, after you have met those agents socially.

systemd is the practical control plane

systemd's own manual is unusually honest: MemoryDenyWriteExecute= is bypassable if the service can write to a non-noexec filesystem (hello /dev/shm) or call memfd_create. The documented countermeasures are not subtle.

For a network-facing unit that does not JIT and does not need anonymous executables:

INI
# /etc/systemd/system/foo.service.d/harden.conf
[Service]
MemoryDenyWriteExecute=yes
InaccessiblePaths=/dev/shm
SystemCallFilter=~memfd_create
NoNewPrivileges=yes
RestrictSUIDSGID=yes
ProtectSystem=strict
PrivateTmp=yes
SystemCallArchitectures=native

Apply this per unit. A global seccomp ban of memfd_create on a Kubernetes node is how you discover, in production, that runc has opinions. Score the result with systemd-analyze security foo.service.

NoExecPaths=/ plus a narrow ExecPaths= is the stricter cousin for a single-binary daemon.

AppArmor, used for what it is good at

Keep the parent service in enforce mode. Deny execution of everything except the service binary and the interpreters it is supposed to have. If a profile would allow /{,usr/}proc/*/fd/* to be executed, stop and ask why. Stock Debian profiles will not carry this technique by themselves. They will, however, make a compromised www-data much less interesting if that service cannot exec a new image.

Close the adjacent doors

A memfd ban that leaves /dev/shm executable is a dress code, not a lock. Hide or noexec the remaining tmpfs surfaces for that unit. shm_open and userland loaders that never call exec are separate problems (mprotect to PROT_EXEC). Landlock can grow teeth for memfd execution on newer kernels; treat that as emerging, not as a Debian 13 default you can cite in a control matrix.

Detection - the ghost has a very obvious ankle bracelet

“Fileless” is a marketing word. The kernel is not shy.

The live indicator that actually works

Shell
for exe in /proc/[0-9]*/exe; do
  target=$(readlink "$exe" 2>/dev/null) || continue
  case "$target" in
    /memfd:*|*memfd:*)
      pid=${exe#/proc/}; pid=${pid%/exe}
      printf '%s %s %s\n' "$pid" "$(tr '\0' ' ' < /proc/$pid/cmdline)" "$target"
      ;;
  esac
done

Complement with r-xp regions in /proc/<pid>/maps backed by anon_inode:[memfd]. lsof will also confess. This is not subtle. It is merely unused on hosts that only scan disks.

auditd, which Debian will not install for you

Plain text
# /etc/audit/rules.d/40-memfd.rules
-a always,exit -F arch=b64 -S memfd_create -F key=anon_file_create
-a always,exit -F arch=b32 -S memfd_create -F key=anon_file_create

-a always,exit -F arch=b64 -S execve,execveat -F key=process_creation
-a always,exit -F arch=b32 -S execve,execveat -F key=process_creation

memfd_create alone is not malice. Programs create non-executable memfds for IPC the way people create temporary buffers. The detection is a sequence: create in pid P, then execve/execveat whose exe= is /memfd:... or whose path is /proc/self/fd/N.

Watch execveat. AT_EMPTY_PATH never needs a pathname, so PATH records may be thin. argv[0] is theatre. exe= is identity.

For low-privilege service accounts, add uid-scoped exec rules rather than only a global firehose. If you rely on auditd, also have a view of ptrace; user-space tampering of the audit daemon is a known sequel, not a plot twist.

Runtime sensors

Falco's proc.is_exe_from_memfd=true is the right shape of rule. Allow-list runc, containerd-shim*, podman, cri-o. Elastic-style process-start queries that match process.executable against memfd: are the same idea with different punctuation.

Without those agents, bpftrace on sys_enter_memfd_create paired with exec probes is a canary, not a SIEM.

A SIEM rule worth keeping:

  1. New process exe matches /memfd: or anon_inode:[memfd].
  2. Parent is not a known container runtime.
  3. User is wrong for that parent (www-data, nobody, an app uid, a fresh account).
  4. Optional: same session recently created an executable memfd, or egress appeared immediately after exec.

False positives, named so they stop being surprises

runc is the one that will humble a naive rule. Newer runc also uses memfd cloning on other paths; Falco has already had to treat this as expected. Allow-list by parent name and exe prefix and cgroup, not by “any memfd.” A user shell whose exe is /memfd: is not runc. Occasional others: some Flatpak/bubblewrap setups, test harnesses, a few language runtimes. Inventory once. Keep the list short.

Incident response when the file was never a file

Live box:

  1. Decide whether to STOP the process. You may lose sockets. Memory capture is kinder to C2 analysis.
  2. Collect /proc/<pid>/{exe,cmdline,environ,cwd,status,maps,fd,net,syscall}.
  3. cp /proc/<pid>/exe /root/recovered.elf often still works. Hash it.
  4. Record parent chain, auid, tty, cgroup, sockets.
  5. Tie the creating pid to audit. Hunt the hash later, knowing disk may never have it.
  6. Ask systemd which unit owns the parent. Ask whether that unit had a seccomp filter or an AppArmor profile that was, in retrospect, decorative.

Dead box: walk fd tables and anonymous inodes in memory. Disk carving is a ritual from a different decade. If you only kept disk backups, you will under-scope the incident. That is a retention problem, not a forensics skill issue.

If the payload is a reverse shell or C2

Here the last layer is almost disappointing, which is the best kind of last layer.

memfd changes where the implant lives. It does not change what an implant must do.

A reverse shell or beacon still needs privileges it already has, sockets it must open, and a process the kernel will list. memfd does not escalate. It does not hide netflow. It does not make ss blush. It does not survive reboot unless something else fetches the image again.

What gets weaker

  • File-hash blocklists.
  • “Wrote an ELF under /tmp, then exec” hunts.
  • AIDE and friends, for the second stage.
  • The comforting sentence “the disk is clean, so we are done.”

What stays honest

Domain Reality
Privileges The implant is the uid that exec’d it.
Network Egress is still egress. Flow logs, conntrack, proxies, DNS.
Process table Visible. Only the exe path is weird.
argv[0] Spoofable. Confuses name hunts. Does not confuse /proc/<pid>/exe.
Credentials Reads of keys, metadata services, and app secrets look like any other stolen process.

A reverse shell is often the noisier of the two: long-lived, one outbound session, children that may be a real /bin/sh (normal exe) under a memfd parent. A beacon is quieter: sleep, jitter, HTTPS that looks like a CDN, no shell until an operator wants one. In both cases the dangerous blend is a parent that already talks to the network. A web worker, a CI runner, a cloud agent. Egress then hides in the service’s baseline.

Persistence is the tell the disk was supposed to be

A memfd image is volatile. Reboot, OOM, crash: gone. That is fine for a session someone wants now. It is unacceptable for an operator who wants Tuesday. So a serious actor still plants a reload path: a unit, a cron, a compromised legitimate binary, a cloud startup script, a container restart policy, a stolen CI job. Those leave control-plane or disk artifacts.

IR translation: killing the process may end the session. If you do not find the reload path, it returns. A host that was only live-exploited may look pristine after reboot. Do not close on “no malware on disk.”

PyLoose is the public object lesson. Cloud Linux workloads, payload decoded in memory, executed from a memfd, miner argv masquerading as something forgettable. No durable second-stage file. Found anyway: exe path, runtime sensors, behavior.

Severity

Do not inflate a ticket because someone said “fileless.” Severity still follows identity, privilege, lateral movement, data access, persistence, and stolen roles. Do inflate uncertainty. A disk-negative malware scan is not exculpatory. Scope includes memory, /proc, audit, and egress from initial access onward. Containment assumes the image can be recopied from the network until the reload path is dead.

The attacker’s bargain, which is a hunting gift

They avoid on-disk scanners and noexec mounts. They spend a distinctive exe path, an unusual syscall pair, and volatility. Many who want more stealth either hide among container runtimes or avoid exec entirely and live in mapped memory. memfd-plus-exec is noisier than a pure in-process loader and quieter than /tmp/.x. If your program of record is file inventory, you will report a clean host while a beacon sits in RAM and talks. If your program of record includes /proc/<pid>/exe and egress, the ankle bracelet is already blinking.

A control order that will not embarrass you

  1. Install auditd. Add memfd_create plus execve/execveat. Ship it to the SIEM.
  2. Set vm.memfd_noexec=1. Investigate failures.
  3. Drop in MDWE, hide /dev/shm, and SystemCallFilter=~memfd_create on internet-facing units that can tolerate it.
  4. Put those units in AppArmor enforce. Unconfined is not a profile; it is a shrug.
  5. Timer or osquery: report /proc/*/exe -> /memfd:.
  6. Consider vm.memfd_noexec=2 only on non-container, non-desktop hosts after a soak.

That stack does not make memfd execution impossible for a privileged, unconstrained process. It makes it rare, logged, and unavailable to the services that are usually the foothold.

The riddle, re-wrapped

Linux can execute a program with no filesystem name because the kernel never required a name. It required a file object. memfd_create is an honest way to get one that dies with its last file descriptor.

The interesting part is not that this is possible. The interesting part is how many defense programs still ask the disk for permission to believe a process exists. The exe symlink has been willing to testify the entire time. The syscalls leave fingerprints. The network still has to go somewhere. The reboot still wipes the stage unless someone built a second act.

The ghost is real. It also files its forwarding address in /proc.

Notes and sources

Kernel and manual material: memfd_create(2), fexecve(3), execveat(2), kernel documentation on non-executable memfds and vm.memfd_noexec, systemd.exec(5) on MemoryDenyWriteExecute.

Runtime and abuse context: runc's sealed-memfd response to CVE-2019-5736; Falco's proc.is_exe_from_memfd rule; Wiz's PyLoose write-up (memfd-executed miner on cloud Linux); MITRE T1620 Reflective Code Loading.

Debian context: AppArmor enabled by default since Buster; Debian 13 tmpfs /tmp; auditd optional; kernel default vm.memfd_noexec=0.

The binary that wasn’t there · Abraxas Labs