Instructions to use touati-kamel/darja-tokenizer-50m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use touati-kamel/darja-tokenizer-50m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="touati-kamel/darja-tokenizer-50m")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("touati-kamel/darja-tokenizer-50m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Darja-Tokenizer-50M: Domain-Aware Byte-Level BPE for Algerian Darija
Darja-Tokenizer-50M: Domain-Aware Byte-Level BPE for Algerian Darija
Darja-Tokenizer-50M is a specialized byte-level Byte-Pair Encoding (BPE) tokenizer trained from scratch on the Algerian Darja Corpus.
It is the official companion tokenizer for Darja-GPT-50M, designed to provide high tokenization efficiency across dialectal Arabic script, Latin-script Arabizi, and French/English code-switching without out-of-vocabulary (OOV) bottlenecks.
Table of Contents
- Why a Custom Darija Tokenizer?
- Fertility Benchmark & Token Efficiency
- Vocabulary & Special Tokens
- Domain Control Tags & Merging Rationale
- Automated Boundary Processing (
TemplateProcessing) - Training Pipeline & Dataset
- Quickstart & Code Examples
- Citation & Contact
Why a Custom Darija Tokenizer?
Standard multilingual tokenizers (such as LLaMA-3, Qwen2.5, or Mistral) are pre-trained predominantly on Modern Standard Arabic (MSA), English, and high-resource languages. When applied to Algerian Darija (الدارجة الجزائرية / Dziri), they fail severely in three ways:
- Catastrophic Fragmentation (Over-Tokenization): Colloquial prefixes (e.g.,
غادي,رانا,ما...ش), dialectal verbal stems, and Maghrebi idioms are shattered into single characters or byte-level pieces, inflating sequence lengths by 50% to 100%. - Context Window Starvation: In a fixed 1,024-token budget, an over-fragmenting tokenizer fits barely ~400 words, choking the effective attention horizon of compact language models.
- Multilingual Code-Switching & Arabizi: Algerian conversations fluidly alternate between Arabic characters and Latin script (Arabizi numbers like
3,7,9and French loanwords). Byte-level BPE natively models raw UTF-8 bytes, guaranteeing zero unknown tokens (<unk>) across arbitrary mixed-script utterances.
Fertility Benchmark & Token Efficiency
Fertility measures the average number of tokens required to represent a single whitespace-delimited word (lower is more efficient).
We evaluated fertility on 335 strictly held-out test documents (~1.2M characters) from the Algerian Darja Corpus that were never seen during tokenizer training:
| Tokenizer | Vocab Size | Tokens / Word (Held-out) | Relative Efficiency | Effective Words per 1,024 Tokens |
|---|---|---|---|---|
| Darja-Tokenizer-50M (Ours) | 24,000 | 1.474 | Baseline (+37.1% win) | ~695 words |
| Qwen/Qwen2.5-0.5B | 151,936 | 2.342 | 37.1% worse | ~437 words |
| Standard LLaMA-3 Tokenizer | 128,256 | 2.610 | 43.5% worse | ~392 words |
Key Takeaway:
Despite having a compact vocabulary of only 24,000 tokens (to keep embedding matrix memory small for 50M-parameter models), our tokenizer represents Algerian Darija with 37.1% fewer tokens than Qwen2.5 (which has a 152K vocabulary).
Vocabulary & Special Tokens
The vocabulary contains 24,000 BPE merge tokens, with dedicated slots reserved for document framing and domain conditioning:
| Token ID | Token | Role | Functional Purpose |
|---|---|---|---|
0 |
<pad> |
Padding token | Batched training and sequence alignment |
1 |
<bos> |
Beginning of Sequence | Prepend marker for document boundaries |
2 |
<eos> |
End of Sequence | Append marker indicating text termination |
3 |
<unk> |
Unknown token | Fallback (virtually unused due to byte-level BPE) |
4 |
<cooking> |
Domain Tag | Culinary recipes, ingredients, food preparation |
5 |
<general> |
Domain Tag | Everyday chats, casual vlogs, comedy, greetings |
6 |
<podcast> |
Domain Tag | Mindset, business, marketing, freelancing |
7 |
<sports> |
Domain Tag | Football match reactions, CAN, leagues, player news |
8 |
<story> |
Domain Tag | Folklore, personal anecdotes, narrative storytelling |
9 |
<tech> |
Domain Tag | Consumer tech, hardware, mobile phones, electronics |
(Tokens 10 through 23,999 correspond to learned byte-level BPE subwords representing frequent Darija morphemes, Arabic words, and Arabizi tokens).
Domain Control Tags & Merging Rationale
During dataset auditing of the 11,151 documents, category volume was analyzed to prevent underrepresented tags from over-memorizing training samples:
DOMAIN_MERGE_MAP = {
"<general>": "<general>",
"<cooking>": "<cooking>",
"<tech>": "<tech>",
"<sports>": "<sports>",
"<podcast>": "<podcast>",
"<story>": "<story>",
"<comedy>": "<general>", # Merged: 11 docs (111.5K words) -> merged into <general>
"<lifestyle_vlog>": "<general>", # Merged: 30 docs (142.7K words) -> merged into <general>
}
- Rationale:
<comedy>and<lifestyle_vlog>had insufficient individual volume (<30 documents). Forcing isolated control tokens would risk verbatim memorization across 6 training epochs. Merging them into<general>concentrated stylistic colloquialisms while keeping the remaining 5 primary domains robustly populated (>200+ docs, 1.5M+ words each).
Automated Boundary Processing (TemplateProcessing)
The tokenizer includes a built-in post-processor that automatically wraps every encoded single string with <bos> and <eos> tokens:
from tokenizers.processors import TemplateProcessing
tokenizer._tokenizer.post_processor = TemplateProcessing(
single="<bos> $A <eos>",
special_tokens=[("<bos>", 1), ("<eos>", 2)],
)
Why this matters:
- In autoregressive causal language modeling, documents are packed into fixed 1,024-token blocks.
- The built-in template ensures that document boundaries are preserved without requiring manual list manipulation before feeding tensors to the model.
Training Pipeline & Dataset
- Corpus Source:
touati-kamel/algerian-darja-corpus(named subsetconditioned) - Training Documents: 10,816 documents (97% split)
- Held-out Documents: 335 documents (3% split, fixed seed
42) - Algorithm: Hugging Face
tokenizers.ByteLevelBPETokenizer - Min Frequency: 2 occurrences
- Format: Exported as Hugging Face
PreTrainedTokenizerFast(providingtokenizer.json,vocab.json,merges.txt,tokenizer_config.json, andspecial_tokens_map.json).
Quickstart & Code Examples
1. Installation
pip install transformers tokenizers
2. Basic Encoding & Decoding
from transformers import AutoTokenizer
TOKENIZER_ID = "touati-kamel/darja-tokenizer-50m"
# Load the fast tokenizer
tokenizer = AutoTokenizer.from_pretrained(TOKENIZER_ID)
text = "اليوم غادي نديرو شربة فريك بنينة بزاف، هايل يا خويا"
tokens = tokenizer.tokenize(text)
input_ids = tokenizer.encode(text)
print(f"Text: {text}")
print(f"Tokens: {tokens}")
print(f"Token IDs: {input_ids}")
# Notice the automatic <bos> (1) at the beginning and <eos> (2) at the end!
3. Handling Mixed Arabic, Arabizi, and French
mixed_text = "salam khoya, chouf had le telephone 3ando batterie kbira w chbab"
print("Tokens:", tokenizer.tokenize(mixed_text))
print("Decoded:", tokenizer.decode(tokenizer.encode(mixed_text)))
4. Domain Conditioning Example
prompt = "<cooking> وصفة طاجين الزيتون على الطريقة الجزائرية"
encoded = tokenizer(prompt, return_tensors="pt")
print("Input IDs:", encoded["input_ids"])
print("First token is <cooking> tag:", tokenizer.decode([encoded["input_ids"][0][1]]))
5. Computing Fertility Against Other Tokenizers
sample_sentence = "البارح تفرجت الماتش تاع ليكيب ناسيونال ولعبو مليح بزاف مع السنغال"
words = len(sample_sentence.split())
darja_tokens = len(tokenizer.encode(sample_sentence, add_special_tokens=False))
print(f"Words: {words}")
print(f"Darja-Tokenizer Tokens: {darja_tokens} (Fertility: {darja_tokens / words:.2f})")
Citation & Contact
If you use Darja-Tokenizer-50M or the Algerian Darja Corpus, please cite:
@misc{touati2026darjatokenizer,
author = {Kamel Touati},
title = {Darja-Tokenizer-50M: Domain-Aware Byte-Level BPE Tokenizer for Algerian Darija},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/touati-kamel/darja-tokenizer-50m}}
}
@misc{touati2026darjacorpus,
author = {Kamel Touati},
title = {Algerian Darja Corpus: Multi-Domain Conversational Transcripts in Algerian Arabic},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus}}
}
Developer: Kamel Touati (Hugging Face Profile)
License: Apache 2.0
Companion Model: Darja-GPT-50M