Cortex-0
FR — Cortex-0 : IA multimodale (texte · image · audio · vidéo) construite entièrement from scratch, sans aucun modèle pré-entraîné ni API externe. EN — Cortex-0: multimodal AI (text · image · audio · video) built entirely from scratch, with no pretrained model and no external API.
Dépôt : Frankenstein-Labs/opendevin-multimodal — renommage « Cortex-0 » différé (ADR-008).
🇫🇷 Français
Statut : ✅ Phases 0 → 12 livrées
Chaque phase a été franchie avec tests verts, commit et publication Hub. Un seul décodeur texte génère ; image, audio et vidéo y sont injectés comme tokens projetés. Tout (tokenizer, attention, RoPE, ViT, encodeur audio, MoE, RAG, API, démo) est codé à la main.
Ce que ce projet EST
- Un transformer texte decoder-only écrit à la main (attention, RoPE, FFN, blocs).
- Un encodeur image ViT maison (patch embedding + blocs), projecteur MLP.
- Un encodeur audio maison sur spectrogrammes mel (mel calculé à la main).
- Un encodeur vidéo : ViT partagé par frame + embedding temporel + transformer spatio-temporel.
- Une fusion multimodale précoce dans une séquence unique de tokens.
- Un MoE (20 experts, top-2) avec routeur et loss d'équilibrage maison.
- Une mémoire RAG maison (embeddings feature-hashing + index cosinus numpy).
- Une API FastAPI, une démo Gradio, un Dockerfile, une publication Hub.
Ce que ce projet N'EST PAS
- ❌ Ni GPT-4V, ni DeepSeek, ni LLaVA. Un modèle de ~15 M de paramètres n'a pas les capacités d'un modèle de plusieurs centaines de milliards. Aucune équivalence n'est prétendue.
- ❌ L'« apprentissage permanent » (Phase 11, mode B) est expérimental et instable : l'oubli catastrophique est démontré par les tests, pas caché.
Contraintes matérielles
Développement, tests et entraînement jouet : CPU uniquement. Tous les chiffres ci-dessous sont mesurés localement sur CPU, jamais inventés.
Résultats mesurés (corpus jouets, CPU)
| Phase | Tâche | Mesure |
|---|---|---|
| 3 | Texte (langage) | perplexité val < 30 atteinte sur corpus jouet |
| 6/8 | Grounding image | exactitude ~0.87–0.92 |
| 7/8 | Grounding audio | exactitude 1.00 |
| 8 | Mixte image+audio | ~0.92 |
| 9 | Vidéo (+ audio) | grounding ~0.78–0.85 (hasard 0.055) |
| 10 | MoE 20 experts | perplexité val 1.005 ; charge H=0.9999 ; 14.86 M params / 5.42 M actifs |
| 12 | Audio via le moteur | transcription correcte sur les voyelles jouets |
Limite honnête : la génération libre d'une description d'image reste faible (accord mot-à-mot parfois partiel) — c'est la limite d'un modèle de cette taille.
Installation
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # puis renseigner HF_TOKEN (jamais en dur)
Utilisation (bibliothèque)
import numpy as np
from src.inference import CortexEngine
eng = CortexEngine(checkpoint="checkpoints/best.pt") # vide => poids init, trained=False
# Texte
print(eng.chat("Bonjour, qui es-tu ?")["text"])
# Image (ndarray uint8, PIL.Image, chemin ou bytes PNG)
img = (np.random.rand(32, 32, 3) * 255).astype("uint8")
print(eng.describe_image(img))
# Audio (échantillons float32, chemin WAV ou bytes WAV)
print(eng.transcribe_audio(np.zeros(8000, dtype="float32")))
# Vidéo (T, H, W, 3) + audio optionnel
frames = np.zeros((4, 32, 32, 3), dtype="uint8")
print(eng.describe_video(frames))
Mémoire / RAG (apprendre un cours)
from src.memory.rag import LessonMemory
mem = LessonMemory()
mem.add_video(
transcript=[(0.0, "Un volcan crache de la lave chaude.")],
frame_captions=[(0.0, "une coulée de lave orange")],
source="volcans.mp4",
)
print(mem.answer("De quoi est faite la lave ?")) # réponse extractive + sources
API REST (FastAPI)
CORTEX_CHECKPOINT=checkpoints/best.pt python -m src.api
# http://localhost:8000/docs
# GET /health /capabilities
# POST /chat {"prompt": "...", "max_new_tokens": 40}
# POST /vision (upload image) POST /audio (upload WAV)
# POST /video (frames .npy) POST /memory/ingest /memory/query
Démo (Gradio)
CORTEX_CHECKPOINT=checkpoints/best.pt python app/gradio_app.py # http://localhost:7860
Docker
docker build -t cortex-0 .
docker run -p 8000:8000 -e CORTEX_CHECKPOINT= cortex-0
Tests
pytest -q # suite complète, tourne sur CPU
ruff check .
Publication Hub
export HF_TOKEN=... # jamais écrit en dur
python -m src.push_to_hub --repo-id Frankenstein-Labs/opendevin-multimodal \
--public --checkpoint checkpoints/best.pt --card README.md \
--message "release: Cortex-0"
Licence
Apache-2.0 — voir LICENSE.
🇬🇧 English
Status: ✅ Phases 0 → 12 shipped
Every phase was crossed with green tests, a commit and a Hub publication. A single text decoder generates; image, audio and video are injected as projected tokens. Everything (tokenizer, attention, RoPE, ViT, audio encoder, MoE, RAG, API, demo) is hand-written.
What this project IS
- A hand-written decoder-only text transformer (attention, RoPE, FFN, blocks).
- A from-scratch image ViT encoder (patch embedding + blocks) + MLP projector.
- A from-scratch audio encoder over mel-spectrograms (mel computed by hand).
- A video encoder: shared per-frame ViT + learned temporal embedding + spatio-temporal transformer.
- Early multimodal fusion into a single token sequence.
- A MoE (20 experts, top-2) with a hand-written router and load-balancing loss.
- A from-scratch RAG memory (feature-hashing embeddings + numpy cosine index).
- A FastAPI service, a Gradio demo, a Dockerfile, a Hub publication.
What this project is NOT
- ❌ Not GPT-4V, DeepSeek or LLaVA. A ~15M-parameter model does not match models with hundreds of billions of parameters. No equivalence is claimed.
- ❌ "Permanent learning" (Phase 11, mode B) is experimental and unstable: catastrophic forgetting is demonstrated by the tests, not hidden.
Hardware constraints
Development, tests and toy training: CPU only. All numbers above are measured locally on CPU, never invented.
Install
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # then set HF_TOKEN (never hard-coded)
Library usage
from src.inference import CortexEngine
eng = CortexEngine(checkpoint="checkpoints/best.pt") # empty => random init, trained=False
print(eng.chat("Hello, who are you?")["text"])
REST API · Demo · Docker · Tests
Same commands as the French section above (python -m src.api, app/gradio_app.py,
docker build, pytest -q).
License
Apache-2.0 — see LICENSE.
Roadmap (phases)
| # | Phase | Statut / Status |
|---|---|---|
| 0 | Fondations / Foundations | ✅ |
| 1 | Tokenizer BPE from scratch | ✅ |
| 2 | Transformer texte / Text transformer | ✅ |
| 3 | Entraînement texte / Text training | ✅ |
| 4 | Encodeur image / Image encoder | ✅ |
| 5 | Alignement image-texte / Image-text alignment | ✅ |
| 6 | VLM génératif / Generative VLM | ✅ |
| 7 | Encodeur audio / Audio encoder | ✅ |
| 8 | Multimodal complet / Full multimodal | ✅ |
| 9 | Vidéo / Video | ✅ |
| 10 | MoE multi-expert (20, top-2) | ✅ |
| 11 | Mémoire / RAG + apprentissage permanent | ✅ |
| 12 | API + Démo + Publication | ✅ |
Détail complet : docs/phases.md.
Choix techniques : docs/architecture.md.
Journal des décisions : docs/decisions.md.