Darja-Tokenizer-50M: Domain-Aware Byte-Level BPE for Algerian Darija

License Dataset Model Vocabulary Size Fertility Win

Darja-Tokenizer-50M is a specialized byte-level Byte-Pair Encoding (BPE) tokenizer trained from scratch on the Algerian Darja Corpus.

It is the official companion tokenizer for Darja-GPT-50M, designed to provide high tokenization efficiency across dialectal Arabic script, Latin-script Arabizi, and French/English code-switching without out-of-vocabulary (OOV) bottlenecks.


Table of Contents

  1. Why a Custom Darija Tokenizer?
  2. Fertility Benchmark & Token Efficiency
  3. Vocabulary & Special Tokens
  4. Domain Control Tags & Merging Rationale
  5. Automated Boundary Processing (TemplateProcessing)
  6. Training Pipeline & Dataset
  7. Quickstart & Code Examples
  8. Citation & Contact

Why a Custom Darija Tokenizer?

Standard multilingual tokenizers (such as LLaMA-3, Qwen2.5, or Mistral) are pre-trained predominantly on Modern Standard Arabic (MSA), English, and high-resource languages. When applied to Algerian Darija (الدارجة الجزائرية / Dziri), they fail severely in three ways:

  1. Catastrophic Fragmentation (Over-Tokenization): Colloquial prefixes (e.g., غادي, رانا, ما...ش), dialectal verbal stems, and Maghrebi idioms are shattered into single characters or byte-level pieces, inflating sequence lengths by 50% to 100%.
  2. Context Window Starvation: In a fixed 1,024-token budget, an over-fragmenting tokenizer fits barely ~400 words, choking the effective attention horizon of compact language models.
  3. Multilingual Code-Switching & Arabizi: Algerian conversations fluidly alternate between Arabic characters and Latin script (Arabizi numbers like 3, 7, 9 and French loanwords). Byte-level BPE natively models raw UTF-8 bytes, guaranteeing zero unknown tokens (<unk>) across arbitrary mixed-script utterances.

Fertility Benchmark & Token Efficiency

Fertility measures the average number of tokens required to represent a single whitespace-delimited word (lower is more efficient).

We evaluated fertility on 335 strictly held-out test documents (~1.2M characters) from the Algerian Darja Corpus that were never seen during tokenizer training:

Tokenizer Vocab Size Tokens / Word (Held-out) Relative Efficiency Effective Words per 1,024 Tokens
Darja-Tokenizer-50M (Ours) 24,000 1.474 Baseline (+37.1% win) ~695 words
Qwen/Qwen2.5-0.5B 151,936 2.342 37.1% worse ~437 words
Standard LLaMA-3 Tokenizer 128,256 2.610 43.5% worse ~392 words

Key Takeaway:

Despite having a compact vocabulary of only 24,000 tokens (to keep embedding matrix memory small for 50M-parameter models), our tokenizer represents Algerian Darija with 37.1% fewer tokens than Qwen2.5 (which has a 152K vocabulary).


Vocabulary & Special Tokens

The vocabulary contains 24,000 BPE merge tokens, with dedicated slots reserved for document framing and domain conditioning:

Token ID Token Role Functional Purpose
0 <pad> Padding token Batched training and sequence alignment
1 <bos> Beginning of Sequence Prepend marker for document boundaries
2 <eos> End of Sequence Append marker indicating text termination
3 <unk> Unknown token Fallback (virtually unused due to byte-level BPE)
4 <cooking> Domain Tag Culinary recipes, ingredients, food preparation
5 <general> Domain Tag Everyday chats, casual vlogs, comedy, greetings
6 <podcast> Domain Tag Mindset, business, marketing, freelancing
7 <sports> Domain Tag Football match reactions, CAN, leagues, player news
8 <story> Domain Tag Folklore, personal anecdotes, narrative storytelling
9 <tech> Domain Tag Consumer tech, hardware, mobile phones, electronics

(Tokens 10 through 23,999 correspond to learned byte-level BPE subwords representing frequent Darija morphemes, Arabic words, and Arabizi tokens).


Domain Control Tags & Merging Rationale

During dataset auditing of the 11,151 documents, category volume was analyzed to prevent underrepresented tags from over-memorizing training samples:

DOMAIN_MERGE_MAP = {
    "<general>": "<general>",
    "<cooking>": "<cooking>",
    "<tech>": "<tech>",
    "<sports>": "<sports>",
    "<podcast>": "<podcast>",
    "<story>": "<story>",
    "<comedy>": "<general>",         # Merged: 11 docs (111.5K words) -> merged into <general>
    "<lifestyle_vlog>": "<general>",  # Merged: 30 docs (142.7K words) -> merged into <general>
}
  • Rationale: <comedy> and <lifestyle_vlog> had insufficient individual volume (<30 documents). Forcing isolated control tokens would risk verbatim memorization across 6 training epochs. Merging them into <general> concentrated stylistic colloquialisms while keeping the remaining 5 primary domains robustly populated (>200+ docs, 1.5M+ words each).

Automated Boundary Processing (TemplateProcessing)

The tokenizer includes a built-in post-processor that automatically wraps every encoded single string with <bos> and <eos> tokens:

from tokenizers.processors import TemplateProcessing

tokenizer._tokenizer.post_processor = TemplateProcessing(
    single="<bos> $A <eos>",
    special_tokens=[("<bos>", 1), ("<eos>", 2)],
)

Why this matters:

  • In autoregressive causal language modeling, documents are packed into fixed 1,024-token blocks.
  • The built-in template ensures that document boundaries are preserved without requiring manual list manipulation before feeding tensors to the model.

Training Pipeline & Dataset

  • Corpus Source: touati-kamel/algerian-darja-corpus (named subset conditioned)
  • Training Documents: 10,816 documents (97% split)
  • Held-out Documents: 335 documents (3% split, fixed seed 42)
  • Algorithm: Hugging Face tokenizers.ByteLevelBPETokenizer
  • Min Frequency: 2 occurrences
  • Format: Exported as Hugging Face PreTrainedTokenizerFast (providing tokenizer.json, vocab.json, merges.txt, tokenizer_config.json, and special_tokens_map.json).

Quickstart & Code Examples

1. Installation

pip install transformers tokenizers

2. Basic Encoding & Decoding

from transformers import AutoTokenizer

TOKENIZER_ID = "touati-kamel/darja-tokenizer-50m"

# Load the fast tokenizer
tokenizer = AutoTokenizer.from_pretrained(TOKENIZER_ID)

text = "اليوم غادي نديرو شربة فريك بنينة بزاف، هايل يا خويا"
tokens = tokenizer.tokenize(text)
input_ids = tokenizer.encode(text)

print(f"Text: {text}")
print(f"Tokens: {tokens}")
print(f"Token IDs: {input_ids}")
# Notice the automatic <bos> (1) at the beginning and <eos> (2) at the end!

3. Handling Mixed Arabic, Arabizi, and French

mixed_text = "salam khoya, chouf had le telephone 3ando batterie kbira w chbab"
print("Tokens:", tokenizer.tokenize(mixed_text))
print("Decoded:", tokenizer.decode(tokenizer.encode(mixed_text)))

4. Domain Conditioning Example

prompt = "<cooking> وصفة طاجين الزيتون على الطريقة الجزائرية"
encoded = tokenizer(prompt, return_tensors="pt")
print("Input IDs:", encoded["input_ids"])
print("First token is <cooking> tag:", tokenizer.decode([encoded["input_ids"][0][1]]))

5. Computing Fertility Against Other Tokenizers

sample_sentence = "البارح تفرجت الماتش تاع ليكيب ناسيونال ولعبو مليح بزاف مع السنغال"

words = len(sample_sentence.split())
darja_tokens = len(tokenizer.encode(sample_sentence, add_special_tokens=False))

print(f"Words: {words}")
print(f"Darja-Tokenizer Tokens: {darja_tokens} (Fertility: {darja_tokens / words:.2f})")

Citation & Contact

If you use Darja-Tokenizer-50M or the Algerian Darja Corpus, please cite:

@misc{touati2026darjatokenizer,
  author       = {Kamel Touati},
  title        = {Darja-Tokenizer-50M: Domain-Aware Byte-Level BPE Tokenizer for Algerian Darija},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/touati-kamel/darja-tokenizer-50m}}
}

@misc{touati2026darjacorpus,
  author       = {Kamel Touati},
  title        = {Algerian Darja Corpus: Multi-Domain Conversational Transcripts in Algerian Arabic},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus}}
}

Developer: Kamel Touati (Hugging Face Profile)
License: Apache 2.0
Companion Model: Darja-GPT-50M

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train touati-kamel/darja-tokenizer-50m