Open Source
Why open source is central to everything we do.
Giving back
The AI community has always moved fastest when knowledge is shared freely. Every Language Matters was built on the shoulders of open source tools, open research, and communities who chose to share rather than hoard. We believe the only credible response to that generosity is to give back — openly, consistently, and without reservation. Every dataset we produce, every tool we build, and every model we train that can be shared, will be shared.
Democratising AI
Proprietary data is one of the most significant barriers to inclusive AI development. When language datasets are locked behind commercial licences, only well-funded organisations can access them — which means the AI built on those datasets continues to serve the already-privileged. By releasing our work under open licences, we ensure that a researcher in Lusaka, a startup in Nairobi, or a university lab in Manila has the same access to high-quality multilingual data as any Silicon Valley company.
Transparency
Open source is also a commitment to accountability. When we publish our datasets, annotation guidelines, and model cards publicly, anyone can inspect our methodology, challenge our decisions, and improve on our work. We do not claim to be perfect — but we do commit to being transparent. Every release includes full documentation of how data was collected, who annotated it, what quality checks were applied, and what limitations exist.
Our open source footprint.
0+
public datasets released
0+
languages represented in open data
0K
downloads across HuggingFace repos
0+
GitHub stars across our repositories
What we publish and release openly.
Everything we release is documented, versioned, and free to use for research and non-commercial applications. Commercial licencing enquiries are welcome.
Annotated datasets
Text classification, NER, sentiment analysis, translation pairs, and transcription datasets across African, South Asian, and Southeast Asian languages. All datasets include full annotation guidelines and quality metadata.
CC BY 4.0Evaluation benchmarks
Standardised evaluation sets for measuring LLM performance on low-resource languages. Designed to enable fair, reproducible comparison of multilingual models on tasks relevant to underrepresented communities.
Apache 2.0Annotation tooling
Open source scripts, schemas, and workflow templates for setting up multilingual annotation pipelines. Includes our quality assurance frameworks and inter-annotator agreement calculators.
MITFine-tuned models
Domain-adapted and language-specific fine-tuned models trained on our annotated corpora, published with full model cards, training details, and evaluation results. Ready to use or fine-tune further.
CC BY-NC 4.0Research & papers
Our research findings, annotation methodology papers, and language documentation reports are published as preprints and open access wherever possible, with accompanying code and data repositories.
Open AccessData schemas & standards
Our general-purpose annotation schema is designed to support translation, summarisation, NER, QA, RAG evaluation, chatbot scoring, and LLM fine-tuning without changing the underlying structure.
MITWhere to find everything we have published.
HuggingFace
huggingface.co/everylanguagematters
Our primary platform for datasets, models, benchmarks, and evaluation resources. Everything is versioned and ready to use via the HuggingFace ecosystem.
GitHub
github.com/everylanguagematters
All tooling, annotation systems, pipelines, and research code are open source. Contributions and discussions are welcome.
How to contribute to our open source work.
You do not need to be a professional engineer to contribute. There are meaningful ways to help at every skill level.
Submit a pull request
Browse open issues on our GitHub repositories and submit fixes, improvements, or new features. All pull requests are reviewed and contributors are credited in release notes.
Review & validate datasets
Native speakers can review published datasets for accuracy, cultural appropriateness, and linguistic quality. Open a GitHub issue or contact us directly to join a validation review.
Improve documentation
Good documentation is what makes open source actually usable. Help us write clearer READMEs, better dataset cards, and more useful tutorials — especially in languages other than English.
Train and share models
If you use our datasets to train models, we encourage you to publish them back to the community on HuggingFace with a model card linking to our data. Tag us so we can amplify your work.
Report issues & suggest languages
Found an error in a dataset? Know a language we should cover next? Open a GitHub issue or email us. Community-driven language prioritisation is how we decide what to work on next.
Share & cite our work
If you use our datasets or tools in your research, please cite us. Visibility in the research community helps us attract contributors, funding, and partnerships that let us do more.
The future of AI is built in the open.
Join researchers, engineers, and linguists from around the world who are building the data infrastructure for an AI that speaks every language. Star our repos, use our datasets, and help us go further — together.
0+
open datasets
0+
languages
0K
downloads
0+
GitHub stars