Nvidia GPU monitoring flaw exposes thousands of servers, risking AI workload disruption
A high-severity vulnerability in Nvidia's DCGM Exporter allowed unauthenticated attackers to crash GPU monitoring on about 2,100 exposed servers, potentially affecting AI workloads.
Security researchers uncovered that thousands of GPU servers were openly exposing Nvidia's DCGM Exporter, which publishes detailed telemetry such as GPU IDs, utilization, and power consumption over plain-text HTTP. The high-severity CVE-2026-47483 vulnerability could be exploited by unauthenticated actors to generate enough requests to deplete memory, causing the exporter to crash and hide GPU health data, thereby disrupting AI training and inference workloads.
Over four scans between March and May, about 2,100 servers across roughly 300 organizations were identified, with 44% of the exposed GPUs situated in the United States and valued at around $100 million, including high-end models like Blackwell Ultra B300, H200, H100, and consumer RTX 5090/4090 cards. Nvidia released a fix in version 4.8.2, and the researchers urged operators to restrict DCGM Exporter, Node Exporter, and Prometheus services to internal monitoring networks.
The study also highlighted public Node Exporter hosts leaking server configurations, which could aid attackers in reconnaissance. Providers such as Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean were notified and have been working with customers to remediate the exposures.
Why it matters
Exposed GPU monitoring can let attackers disrupt AI workloads and reveal sensitive infrastructure details.
In this story
