CVE Catalog

CVE-2026-89812

Unknown
Published: Updated: Translated: NVD NIST

Summary

In the Linux kernel DRM/AMDGPU subsystem, MES ring fences are not force-completed during reset. After a reset, the first MES submission may poll forever, causing resume failure and system hang.

Risk Assessment

The risk includes failed resume after a GPU reset, potentially requiring a reboot and leading to data loss.

Recommendation

It is recommended to update the Linux kernel to a version with the fix that force-completes MES ring fences for all XCCs.

Other vulnerabilities in Linux kernel

See all
Original NVD description (English source)

In the Linux kernel, the following vulnerability has been resolved: drm/amdgpu: force complete the MES ring fences on reset The MES scheduler ring has no drm scheduler (no_scheduler = true), so it is skipped by the force-completion loop in amdgpu_device_pre_asic_reset(). It uses a polling fence whose hw value lives in wb (GTT) memory and survives a MODE1 reset, while fence_drv.sync_seq keeps advancing for every packet. When the reset is triggered because MES itself stopped responding, the timed-out packets advance sync_seq past the last hw fence value MES wrote. After resume the first MES submission polls forever on a seq that is never written back, failing the resume and wedging the box on a second reset: amdgpu: MES ring buffer is full. amdgpu: *ERROR* ring gfx_0.0.0 test failed (-110) amdgpu: resume of IP block <gfx_v11_0> failed -110 amdgpu: GPU reset end with ret = -110 Force complete the MES scheduler ring fences together with the scheduler rings so their hw fence is realigned to sync_seq. v2: cover all XCCs (one scheduler ring each), not just mes.ring[0].

Vulnerability data from NVD (NIST) · CISA KEV · EPSS