CVE-2026-24206 is an authentication-bypass vulnerability in NVIDIA Triton Inference Server versions before 26.03. The flaw is classified as CWE-288 (Authentication Bypass Using an Alternate Path or Channel). It permits an unauthenticated remote attacker to bypass intended authentication controls under affected deployments.
Mallory correlates every CVE against your assets, your vendors, and active adversary campaigns. Know which vulnerabilities matter for you, not just which ones are loud.
What it means. What to do now. Patch path, mitigations, and the assume-compromise checklist.
What an attacker gets, and what they’ve been doing with it.
If you can’t patch tonight, do this now.
Patch, then assume compromise.
2 valid exploits after Mallory filtered fakes, detection scripts, and README-only repos.
Repository contains a standalone Python proof-of-concept and supporting documentation for CVE-2026-24207, an authentication bypass affecting NVIDIA Triton Inference Server SageMaker HTTP model-management endpoints, with analysis of the structurally similar CVE-2026-24206 for Vertex AI. The repo has 13 files total, with 5 code/config-like files: exploit.py, detect.py, bypass_demo.py, example/model/1/model.py, and example/model/config.pbtxt. Primary entry points are the three Python scripts. Main exploit capability is in exploit.py. In default probe mode it sends unauthenticated requests to the SageMaker /models surface to demonstrate that GET /models, POST /models, and DELETE /models/<name> reach the handler without the operator’s configured restricted header. In rce mode it sends a POST /models request with JSON containing model_name and a local filesystem url, plus the X-Amzn-SageMaker-Target-Model header, to force Triton to load a model from a target-local path. If that path contains a valid Python-backend Triton model, model.py executes at import time and again during TritonPythonModel.initialize(), yielding code execution as the Triton process user. The repository also includes detect.py, which is a non-destructive vulnerability checker that probes GET /models without auth and classifies targets as PATCHED, VULNERABLE, or UNKNOWN based on HTTP status and response text. bypass_demo.py is an educational script showing side-by-side behavior for /ping, /models with the expected auth header, and /models without it, illustrating the bypass mechanics without changing state. The example/ directory provides a benign payload model used for RCE demonstration. example/model/config.pbtxt defines a minimal Triton Python backend model, and example/model/1/model.py writes marker files to /tmp to prove code execution. This is not a stealth payload; it is a basic operational demonstration payload. Supporting docs explain root cause, patch diff, and the RCE chain. detection/suricata.rules and detection/triton-access.md provide defender-focused detection content, including version fingerprinting via GET /v2 on port 8000 and behavioral detection of bursts of /models requests on port 8080. Overall, this is a real exploit repository, not just a detector. It demonstrates unauthenticated access to restricted model-management APIs and, when the filesystem prerequisite is met, pre-auth remote code execution through Triton’s Python backend model loading path.
Repository contains a complete Python proof-of-concept and impact demonstration for NVIDIA Triton Inference Server SageMaker auth bypass CVE-2026-24207, with analysis of the structurally similar Vertex AI issue CVE-2026-24206. The main exploit file is exploit.py, which supports two modes: probe mode sends unauthenticated GET /models, POST /models, and DELETE /models/<name> requests to show that model-management handlers are reachable without the configured restricted header; rce mode sends an unauthenticated POST /models with JSON body {model_name, url} and X-Amzn-SageMaker-Target-Model header to force Triton to load a model from a target-local filesystem path. If that path contains a valid Python-backend Triton model, model.py executes as the Triton process user, yielding pre-auth RCE. The included example/model payload is benign and writes marker files under /tmp to prove execution. Additional scripts include detect.py, a read-only vulnerability checker that classifies PATCHED/VULNERABLE/UNKNOWN based on GET /models behavior, and bypass_demo.py, an educational script comparing authorized and unauthorized access to /models. Supporting documentation explains root cause, patch diff, and the RCE chain, while detection/ contains Suricata rules and operational detection guidance. Overall, this is a real exploit repository with both safe detection tooling and an operational unauthenticated RCE demonstration contingent on a readable attacker-staged model path.
Products and vendors Mallory has correlated with this vulnerability. Open in Mallory to drill down to specific CPE configurations and version ranges.
Vendor-confirmed product mapping. Mallory continuously reconciles this list against your asset inventory.
6 sources tracked across advisories, community write-ups, and news. New activity surfaces here as Mallory finds it.
Unknown
An authentication bypass vulnerability in NVIDIA Triton Inference Server that could lead to privilege escalation, denial of service, or information disclosure.
Query your assets running an affected version, and investigate the blast radius.
Every observed campaign linking this CVE to a named adversary.
Malware families riding this exploit, with evidence and IOCs.
YARA, Sigma, Snort, and vendor rules, auto-deployed to your SIEM.
Cross-references every affected SKU, including bundled OEM variants.
Community discussion across Reddit, Mastodon, and other social sources.