
Addressing the AI Accuracy Gap: Open-Telco LLM Benchmarks
The GSMA is spearheading a vital, global, open-source initiative. Our mission: To develop specialised benchmarks, datasets, and models that empower large language models to truly understand and serve the unique, complex needs of the telecommunications industry.
Today’s LLMs don’t speak telco. And that’s a problem.
General-purpose LLMs demonstrate impressive capabilities in many domains. However, they consistently falter when confronted with the intricate and highly specialised knowledge required by the telecommunications ecosystem. This gap isn’t just a minor issue; it’s a critical barrier to AI-driven innovation and efficiency for telcos.
Poor accuracy on technical queries
Tests show generic LLMs (like GPT-4 & Claude) can provide 30-40% incorrect responses when quizzed on specific technical topics. Queries like ‘RAN slicing vs. Network slicing’ often confuse crucial 3GPP standards.
Hallucinations in regulations and standards
When asked about vital areas like 5G spectrum policies, AI has been found to invent non-existent frequency bands and misquote ITU rules. This poses significant regulatory risks.
Inadequate for network troubleshooting
Real-world trials have seen AI suggest ‘fixes’ that would actually worsen issues like packet loss, critically failing to interpret vendor-specific network KPIs (e.g. SNR, RRC success rate).
Why this matters
For telcos, this translates into wasted investments in AI solutions that don’t deliver. It also means increased regulatory exposure, compromised user experiences, and the strategic risk of falling behind more agile, AI-native competitors.
Our core goals
- Unite industry and boost performance: Foster a collaborative ecosystem to collectively enhance LLM performance for specific telecommunications applications.
- Promote holistic evaluation: Drive comprehensive benchmarking that considers not just telco capabilities, but also energy efficiency, safety, and multilingual support.
- Accelerate targeted innovation: Provide the essential tools, datasets, and frameworks needed to develop telco-specialised AI models that deliver tangible value to the telecom sector and deploy them.
How we’re achieving this
- Use case driven: Telecom operators submit real-world use cases and critical requirements to the GSMA, ensuring benchmarks address genuine industry needs.
- Community built: The Open-Telco Team, comprised of experts from across the industry, creates robust, open-source LLM benchmarks based on these prioritized use cases.
- Transparent evaluation and open access: GSMA facilitates the running of these benchmarks on a variety of telco-specific and general-purpose LLMs, publishing the results openly.
(Version 1 is now live on Hugging Face!)
Your expertise. Our platform. Transformative impact.
The Open-Telco LLM Benchmarks initiative is more than a technical project; it’s a strategic imperative to ensure AI effectively serves the telecommunications industry. Your participation offers direct benefits and the chance to co-create the future.
Join the vanguard of telco AI. Your contribution matters.
Ready to help shape the future of AI in telecoms?
We invite all passionate stakeholders – telecom operators, AI developers, academic institutions, and open-source enthusiasts – to lend their expertise to the Open-Telco LLM Benchmarks initiative. There are numerous impactful ways to contribute:
- Telecom operators: Submit LLM use cases and define requirements.
- Academic and research partners: Assist with feasibility analysis, open data review, and benchmark definition; help build playbooks.
- AI model developers and open-source community: Lend your expertise, contribute models for benchmarking, help build models for use cases, and submit results to the leaderboard.
How to participate
Submit use cases
Email aiusecase@gsma.com (your submissions will be aggregated and anonymised).
Express interest and general inquiries:
Contact us at lpowell@gsma.com to discuss how you can contribute your expertise, join sub-groups, or learn more.
Or, register your interest via our form:



