Pandex researchers found that stale or invalid package and domain references in vendor-hosted llms.txt and llms-full.txt files can turn AI coding agents into a software supply-chain entry point. Across 8,565 files, they identified 237 references to nonexistent, mistyped, relocated, expired, or outdated packages and domains; a separate review of 120 sites found 227 unsafe instructions. By registering available Python and npm package names and using benign callback proof-of-concept packages, the researchers observed agent-driven installations contacting their infrastructure in as little as four minutes.
The tests implicated agents associated with Claude, OpenAI Codex, and Nous Research Hermes, with more autonomous models showing markedly higher execution rates; GPT-5 Luna and Sol reportedly exceeded 90%, while Claude Opus 4.8 at medium effort executed the test payload roughly 30% of the time. The risk has already been illustrated by Clerk documentation that directed users to install the nonexistent npm package clerk-next-fix-auth-protection; an unknown actor claimed that name and used it to distribute malware. Organizations should treat AI-readable external documentation as untrusted input and require package ownership, namespace, domain, and repository verification before agents run installation commands, since EDR tools may classify npm or PyPI activity as normal development behavior.

Trace attribution and downstream blast radius.
5 events from the most recent confirmed update back to the earliest known activity.
After researchers notified Clerk about the unsafe reference, Clerk fixed the issue. Clerk said installation of the relevant binary from @clerk/eslint-plugin was not dangerous, but acknowledged that other cases could have resulted in malware installation.
An llms.txt file on Clerk's site included an npx command for the nonexistent clerk-next-fix-auth-protection package. An unknown actor registered that npm name and uploaded malware, meaning an AI agent following the official-site instruction could download and execute malicious code.
Pandex reported that more autonomous frontier models were more likely to install and execute the proof-of-concept malware when following malicious llms.txt guidance. GPT-5 Luna and Sol reportedly executed the payload at rates above 90%, while Claude Opus 4.8 at medium effort did so about 30% of the time.
The researchers registered available package names and deployed Python and Node.js proof-of-concept packages that called back to researcher-controlled infrastructure when executed. They received agent-linked callbacks shortly after the packages became available, including from organizations and companies using Claude, OpenAI Codex, and Nous Research Hermes.
Researchers reviewing thousands of llms.txt and llms-full.txt files found stale, nonexistent, mistyped, relocated, expired, or otherwise unsafe references to software packages and domains. Some files included pip, npm, or other installation commands that could direct an AI agent to attacker-claimed resources.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
See attribution and downstream blast radius, and whether this package or vendor reaches your builds.
3 references tracked. Mallory keeps watching after this page renders.
schneier.com
Open sourcetomshardware.com
Open sourcexakep.ru
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.