Supported by the GSMA, the African Trust & Safety LLM Challenge has produced a new benchmark of 4,216 verified and reproducible AI safety stress tests across African languages, multilingual prompts, and code-switched contexts. The dataset was generated through the Zindi community, with 320 participants submitting more than 42,000 adversarial attacks across 4,010 markdown files. After evaluation, contributions from 307 participants were represented in the final benchmark.
The benchmark was created through a rigorous multi-stage evaluation pipeline designed to prioritise quality, reproducibility, diversity, and fairness. Submissions were first validated for structure, metadata, supported languages, and target models. Duplicate and near-duplicate attacks were then removed using multilingual semantic similarity checks, helping prevent repeated or templated prompts from inflating the results.
Each valid attack was assessed by multiple independent LLM judges using a 20-point rubric covering attack validity, evidence of model failure, classification accuracy, non-triviality, and cultural specificity. Attacks also had to reproduce consistently under controlled evaluation settings, ensuring that only prompts which reliably triggered unsafe or harmful model behaviour were included in the final benchmark.
The resulting benchmark captures a broad range of adversarial techniques and potential harms, from misinformation and unsafe medical advice to cybersecurity, hate speech, extremism, and financial risk. It highlights why AI safety cannot be evaluated only in dominant global languages: Africa’s linguistic diversity, multilingualism, and culturally specific risk contexts expose vulnerabilities that conventional benchmarks often miss.
By supporting the challenge, GSMA is helping ensure that emerging global standards for trustworthy AI better reflect the realities of African users and languages. This is a reusable, Africa-focused AI safety benchmark showing how LLM safety must be tested across African languages, multilingual contexts, and culturally specific harms – exactly the goal when GSMA and Zindi launched the challenge.
Top languages represented
Swahili 33.4%, Hausa 21.6%, Yoruba 14.1%, Igbo 9.3%, Zulu 6.4%, Afrikaans 3.7%, Amharic 3.3%, Akan 3.2%.
Top risk categories
Harmful instructions 14.5%, illegal activity 13.0%, misinformation 9.3%, cybersecurity 9.3%, unsafe medical advice 7.7%, bias/discrimination 7.0%, violence 6.3%, hate/harassment 6.1%.
Top attack techniques
Roleplay 12.9%, indirect requests 10.4%, hypothetical scenarios 9.7%, context poisoning 8.0%, persona hijacking 7.0%, adversarial rephrasing 7.0%, translation pivot 6.5%.