CVE-2026-93241
Low risk· EPSS 12%Exploitation Probability (EPSS)
Low risk12th percentile - higher than 12% of all known CVEs
Summary
A vulnerability in the Linux kernel's memory cgroup (memcg) mechanism causes OOM-killed processes to get stuck in the exit path for hours. The issue occurs when the oom_reaper fails to free the process memory (setting MMF_OOM_SKIP), and dying threads still attempt to swap in pages from zswap, leading to serialization on oom_lock and severe exit delays.
Risk Assessment
Organizations may face prolonged application outages where OOM-killed processes do not terminate, holding system resources and requiring manual intervention. In extreme cases, processes can hang for many hours, causing service unavailability.
Recommendation
Apply the Linux kernel patch that bypasses reclaim and the OOM killer for dying tasks once the oom_reaper is done. Update the kernel to a version containing this fix and monitor systems for similar symptoms.
Other vulnerabilities in Linux kernel
See all- CVE-2026-98162Unknown
In the Linux kernel, the smb2_tree_connect() function of the SMB server (ksmbd) leaks a tree connection. When ksmbd_iov_pin_rsp() fails, the newly created tree connection is not disconnected, leading to a resource leak.
- CVE-2026-98160Unknown
In the Linux kernel, the staging rtl8723bs driver's rtw_sdio_if1_init() function frees padapter->HalData with kfree(), even though it was allocated via vzalloc(). Using kfree() to release a vmalloc-backed buffer can lead to memory corruption.
- CVE-2026-100079Unknown
In the Linux kernel, the USB Type-C (ucsi) subsystem's ucsi_register() creates per-instance debugfs entries, but ucsi_unregister() keeps them until ucsi_destroy(). Drivers like ucsi_glink that unregister/register the same UCSI instance across remoteproc restart then try to create an already existing debugfs directory.
- CVE-2026-100078Unknown
In the Linux kernel, the iwlwifi (mei) driver's iwl_mei_write_cyclic_buf() function receives an incorrect first argument — the q_head pointer is passed instead of cldev. The bug has been fixed.
- CVE-2026-100077Unknown
In the Linux kernel, the drm/msm driver does not safely retire a hung submit before GPU recovery completes. Retiring the submit triggers BO free, which can result in GPU pagefaults since the GPU may be actively accessing those BOs.
- CVE-2026-100074Unknown
In the Linux kernel, a bug was fixed where the BPF_REFCOUNT field was not marked as unique, although it should be. The fix addresses this oversight.
- CVE-2026-100073Unknown
In the Linux kernel, a bug in the ext4 filesystem related to transaction overflow during writeback was fixed. A previous fix was too eager in reducing reserved transaction credits, leading to insufficient reservation in some corner cases. The fix uses ext4_meta_trans_blocks() for a correct upper bound estimate.
- CVE-2026-100072Unknown
In the Linux kernel, a problem in the ACPI subsystem was fixed where the use of acpi_get_first_physical_node() in acpi_platform_fill_resource() and acpi_create_platform_device() was unsafe because the returned device could be freed at any time. The fix replaces it with acpi_bus_get_primary_device() and adjusts the code to call it only once.
- CVE-2026-100071Unknown
In the Linux kernel, a memory leak in the HSR (High-availability Seamless Redundancy) module was fixed. When hsr_dev_finalize() fails after registering an RX handler, dynamic nodes learned in that window are not released. The fix frees both dynamic databases in the error unwind path.
- CVE-2026-100070Unknown
In the Linux kernel, a bug in the netfilter nf_nat_sip module was fixed where the offset was not rewound when NAT shrinks the packet. If map_addr() changes the packet length, coff may point to an incorrect position, potentially causing subsequent Contact headers to be skipped and leaking internal network details.
Original NVD description (English source)
In the Linux kernel, the following vulnerability has been resolved: memcg: bypass the reclaim and oom killer for dying tasks once oom_reaper is done At Meta, we are seeing instances where an OOM killed job is stuck in the exit path for several hours. In one particular case, the job was stuck for more than 8 hours and I had to manually remove the memory.max limits to allow the process to exit. The job was a single process job and had ~55 GiB memory.max and zswap enabled. It had almost 0 anon in memory and ~111 GiB in zswap compressed to ~51 GiB zswap pool (i.e. almost all of memory.current was zswap). Nothing was left on the LRUs to reclaim. On further inspection, I observed ~20k threads of that process stuck with the following stack: [<0>] mem_cgroup_out_of_memory+0x4e/0xa0 [<0>] charge_memcg+0x8bf/0x990 [<0>] mem_cgroup_swapin_charge_folio+0x4e/0x80 [<0>] __read_swap_cache_async+0x10c/0x260 [<0>] swapin_readahead+0x116/0x3f0 [<0>] do_swap_page+0x13c/0x1ce0 [<0>] handle_mm_fault+0x61d/0x11f0 [<0>] do_user_addr_fault+0x3e7/0x6d0 [<0>] exc_page_fault+0x8f/0x110 [<0>] asm_exc_page_fault+0x22/0x30 [<0>] __get_user_8+0x14/0x20 [<0>] futex_cleanup+0x27/0x1c0 [<0>] futex_exit_release+0x47/0x60 [<0>] do_exit+0x107/0x940 [<0>] do_group_exit+0x81/0xa0 [<0>] get_signal+0x2b1/0x6e0 [<0>] arch_do_signal_or_restart+0x1a/0x1c0 [<0>] exit_to_user_mode_loop+0xa8/0x1c0 [<0>] do_syscall_64+0x152/0x250 [<0>] entry_SYSCALL_64_after_hwframe+0x4b/0x53 In addition the dmesg was filled with "Out of memory and no killable processes..." messages. I have no idea why oom reaper was not able to reap/unmap the process. My guess is that since oom reaper tries to acquire mmap_lock in read mode limited number of times and then gives up, there might be a thread of that process which had mmap_lock in write mode at that time. My initial suspicion was the futex_cleanup and kernel page fault causing infinite fault and charge retries but that was put to rest in previous discussions happened on similar problem [1]. My current theory is that it is just a simple slow serialization behind the oom_lock. Unlike page allocator, memcg charge code takes the oom_lock without the "try". Though memcg oom code uses mutex_lock_killable(), note that in the call stack get_signal() consumes SIGKILL (or sigdelset(SIGKILL)) before calling do_group_exit(). So this mutex_lock_killable() is just a mutex_lock() here. Therefore 10s of thousands of threads are waiting on oom_lock and one by one they get -EFAULT from get_user() in the futex cleanup code and bails out. Discussion from [1] led to commit a75ffa26122b ("memcg, oom: do not bypass oom killer for dying tasks") which routes dying tasks into the OOM path precisely so the oom_reaper can reap their mm and free the memory asynchronously. But the reaper is best-effort and one-shot: if it cannot take mmap_lock for read (e.g. a sibling thread holds it for write) it sets MMF_OOM_SKIP and never retries, leaving only the glacial oom_lock-serialized synchronous drain. Once MMF_OOM_SKIP is set there is no more asynchronous reclaim coming for the mm, so a dying task charging against it has nothing left to wait for: it frees its memory only once it finishes exiting. Running reclaim and the (no-victim) OOM killer for it is then pointless, and doing it for 10s of thousands of exiting threads is what serializes them behind oom_lock. So before reclaim, if current is an OOM victim whose reaper is done, fail the charge. Reproduced with 20k threads, each parking a robust futex head on its own zswapped page, OOM-group-killed while a sibling holds mmap_lock for write so the reaper gives up and sets MMF_OOM_SKIP. Tested on next-20260728 and baseline show ~90 seconds exit time while with the patch the exit time reduced to ~3 seconds.
Vulnerability data from NVD (NIST) · CISA KEV · EPSS

