Two separate items highlight operational security risks created by poorly governed AI systems: (1) an AI chatbot that discloses sensitive “flags” when directly asked, and (2) a vulnerability disclosure attempt that appears to be blocked by an AI-managed inbox that does not reliably route reports to humans. Both point to a common failure mode where AI agents are deployed without adequate guardrails, escalation paths, or verification that security-critical communications reach accountable staff.
A TryHackMe write-up describes an intentionally “easy” CupidBot room where the bot returns the requested flags simply when prompted, underscoring how prompt/role-play interactions can bypass intended controls when the model is not hardened against data exfiltration requests. Separately, an Ask HN post describes a “vibe coded” app exposing user chats and identity details; the reporter says email contact is answered by AI promising internal escalation, but no fix or human response follows, raising the risk that AI front-ends can silently fail responsible disclosure workflows and prolong exposure of sensitive user data.

Track how attackers are adapting to this technology.
3 events from the most recent confirmed update back to the earliest known activity.
A write-up on the TryHackMe "CupidBot" room reported that the chatbot would reveal flags when directly asked, without requiring any command injection. The author characterized this as a prompt-hardening failure rather than a true command-injection vulnerability.
The reporter attempted to disclose the exposure by email, but received only an automated AI reply saying the issue would be raised internally. They also messaged the team on X and got no response, leaving uncertainty about whether any human had seen the report.
A reporter found that a "vibe coded" application was exposing user chat content and user identity details because of an apparent implementation mistake rather than an exploit. The issue appeared obvious enough that the reporter suspected others may also have noticed it.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
infosecwriteups.com
Open sourcenews.ycombinator.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.