Researchers released XL-SafetyBench, a country-grounded evaluation suite designed to measure LLM jailbreak resistance and cultural sensitivity in native-language contexts rather than translated English-centric tests. The benchmark contains 5,500 cases across 10 country-language pairs: 4,500 adversarial jailbreak prompts and 1,000 benign prompts involving culturally specific sensitivities. It uses Attack Success Rate (ASR), Cultural Sensitivity Rate (CSR), and the new Neutral-Safe Rate (NSR) metric to separate legitimate safety refusals from irrelevant or degraded outputs caused by weak language capability.
Testing of 10 frontier and 27 local models found large differences in resilience and localized safety performance. Claude-4.5-Sonnet recorded a 2.8% average jailbreak ASR, compared with 98.8% for Mistral-Large-3, while frontier models averaged 49.9% on cultural sensitivity. Frontier-model jailbreak robustness did not reliably correlate with cultural awareness; among local models, an inverse relationship between ASR and NSR indicated that seemingly safe outcomes can result from generation failure rather than genuine alignment. The findings highlight the need for region-specific red teaming and evaluation before deploying LLMs across multilingual markets.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
Dasol Choi and 16 co-authors submitted the XL-SafetyBench paper, introducing a country-grounded benchmark for evaluating LLM jailbreak robustness and cultural sensitivity. The benchmark contains 5,500 native-language test cases across 10 country-language pairs and evaluates ASR, NSR, and CSR.
AIM Intelligence published the XL-SafetyBench dataset for evaluating LLM jailbreak robustness and cultural sensitivity across 10 country-language pairs. The release includes 4,500 adversarial jailbreak prompts, 1,000 cultural-sensitivity scenarios, and Parquet dataset files with evaluation metrics.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
4 references tracked. Mallory keeps watching after this page renders.
aim-intelligence.com
Open sourcearxiv.org
Open sourcegithub.com
Open sourcehuggingface.co
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.