Cortex-0

FR — Cortex-0 : IA multimodale (texte · image · audio · vidéo) construite entièrement from scratch, sans aucun modèle pré-entraîné ni API externe. EN — Cortex-0: multimodal AI (text · image · audio · video) built entirely from scratch, with no pretrained model and no external API.

Dépôt : Frankenstein-Labs/opendevin-multimodal — renommage « Cortex-0 » différé (ADR-008).


🇫🇷 Français

Statut : ✅ Phases 0 → 12 livrées

Chaque phase a été franchie avec tests verts, commit et publication Hub. Un seul décodeur texte génère ; image, audio et vidéo y sont injectés comme tokens projetés. Tout (tokenizer, attention, RoPE, ViT, encodeur audio, MoE, RAG, API, démo) est codé à la main.

Ce que ce projet EST

  • Un transformer texte decoder-only écrit à la main (attention, RoPE, FFN, blocs).
  • Un encodeur image ViT maison (patch embedding + blocs), projecteur MLP.
  • Un encodeur audio maison sur spectrogrammes mel (mel calculé à la main).
  • Un encodeur vidéo : ViT partagé par frame + embedding temporel + transformer spatio-temporel.
  • Une fusion multimodale précoce dans une séquence unique de tokens.
  • Un MoE (20 experts, top-2) avec routeur et loss d'équilibrage maison.
  • Une mémoire RAG maison (embeddings feature-hashing + index cosinus numpy).
  • Une API FastAPI, une démo Gradio, un Dockerfile, une publication Hub.

Ce que ce projet N'EST PAS

  • ❌ Ni GPT-4V, ni DeepSeek, ni LLaVA. Un modèle de ~15 M de paramètres n'a pas les capacités d'un modèle de plusieurs centaines de milliards. Aucune équivalence n'est prétendue.
  • ❌ L'« apprentissage permanent » (Phase 11, mode B) est expérimental et instable : l'oubli catastrophique est démontré par les tests, pas caché.

Contraintes matérielles

Développement, tests et entraînement jouet : CPU uniquement. Tous les chiffres ci-dessous sont mesurés localement sur CPU, jamais inventés.

Résultats mesurés (corpus jouets, CPU)

Phase Tâche Mesure
3 Texte (langage) perplexité val < 30 atteinte sur corpus jouet
6/8 Grounding image exactitude ~0.87–0.92
7/8 Grounding audio exactitude 1.00
8 Mixte image+audio ~0.92
9 Vidéo (+ audio) grounding ~0.78–0.85 (hasard 0.055)
10 MoE 20 experts perplexité val 1.005 ; charge H=0.9999 ; 14.86 M params / 5.42 M actifs
12 Audio via le moteur transcription correcte sur les voyelles jouets

Limite honnête : la génération libre d'une description d'image reste faible (accord mot-à-mot parfois partiel) — c'est la limite d'un modèle de cette taille.

Installation

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env   # puis renseigner HF_TOKEN (jamais en dur)

Utilisation (bibliothèque)

import numpy as np
from src.inference import CortexEngine

eng = CortexEngine(checkpoint="checkpoints/best.pt")  # vide => poids init, trained=False

# Texte
print(eng.chat("Bonjour, qui es-tu ?")["text"])

# Image (ndarray uint8, PIL.Image, chemin ou bytes PNG)
img = (np.random.rand(32, 32, 3) * 255).astype("uint8")
print(eng.describe_image(img))

# Audio (échantillons float32, chemin WAV ou bytes WAV)
print(eng.transcribe_audio(np.zeros(8000, dtype="float32")))

# Vidéo (T, H, W, 3) + audio optionnel
frames = np.zeros((4, 32, 32, 3), dtype="uint8")
print(eng.describe_video(frames))

Mémoire / RAG (apprendre un cours)

from src.memory.rag import LessonMemory
mem = LessonMemory()
mem.add_video(
    transcript=[(0.0, "Un volcan crache de la lave chaude.")],
    frame_captions=[(0.0, "une coulée de lave orange")],
    source="volcans.mp4",
)
print(mem.answer("De quoi est faite la lave ?"))   # réponse extractive + sources

API REST (FastAPI)

CORTEX_CHECKPOINT=checkpoints/best.pt python -m src.api
# http://localhost:8000/docs
# GET  /health   /capabilities
# POST /chat     {"prompt": "...", "max_new_tokens": 40}
# POST /vision   (upload image)      POST /audio  (upload WAV)
# POST /video    (frames .npy)       POST /memory/ingest  /memory/query

Démo (Gradio)

CORTEX_CHECKPOINT=checkpoints/best.pt python app/gradio_app.py   # http://localhost:7860

Docker

docker build -t cortex-0 .
docker run -p 8000:8000 -e CORTEX_CHECKPOINT= cortex-0

Tests

pytest -q     # suite complète, tourne sur CPU
ruff check .

Publication Hub

export HF_TOKEN=...   # jamais écrit en dur
python -m src.push_to_hub --repo-id Frankenstein-Labs/opendevin-multimodal \
    --public --checkpoint checkpoints/best.pt --card README.md \
    --message "release: Cortex-0"

Licence

Apache-2.0 — voir LICENSE.


🇬🇧 English

Status: ✅ Phases 0 → 12 shipped

Every phase was crossed with green tests, a commit and a Hub publication. A single text decoder generates; image, audio and video are injected as projected tokens. Everything (tokenizer, attention, RoPE, ViT, audio encoder, MoE, RAG, API, demo) is hand-written.

What this project IS

  • A hand-written decoder-only text transformer (attention, RoPE, FFN, blocks).
  • A from-scratch image ViT encoder (patch embedding + blocks) + MLP projector.
  • A from-scratch audio encoder over mel-spectrograms (mel computed by hand).
  • A video encoder: shared per-frame ViT + learned temporal embedding + spatio-temporal transformer.
  • Early multimodal fusion into a single token sequence.
  • A MoE (20 experts, top-2) with a hand-written router and load-balancing loss.
  • A from-scratch RAG memory (feature-hashing embeddings + numpy cosine index).
  • A FastAPI service, a Gradio demo, a Dockerfile, a Hub publication.

What this project is NOT

  • ❌ Not GPT-4V, DeepSeek or LLaVA. A ~15M-parameter model does not match models with hundreds of billions of parameters. No equivalence is claimed.
  • ❌ "Permanent learning" (Phase 11, mode B) is experimental and unstable: catastrophic forgetting is demonstrated by the tests, not hidden.

Hardware constraints

Development, tests and toy training: CPU only. All numbers above are measured locally on CPU, never invented.

Install

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env   # then set HF_TOKEN (never hard-coded)

Library usage

from src.inference import CortexEngine
eng = CortexEngine(checkpoint="checkpoints/best.pt")   # empty => random init, trained=False
print(eng.chat("Hello, who are you?")["text"])

REST API · Demo · Docker · Tests

Same commands as the French section above (python -m src.api, app/gradio_app.py, docker build, pytest -q).

License

Apache-2.0 — see LICENSE.


Roadmap (phases)

# Phase Statut / Status
0 Fondations / Foundations ✅
1 Tokenizer BPE from scratch ✅
2 Transformer texte / Text transformer ✅
3 Entraînement texte / Text training ✅
4 Encodeur image / Image encoder ✅
5 Alignement image-texte / Image-text alignment ✅
6 VLM génératif / Generative VLM ✅
7 Encodeur audio / Audio encoder ✅
8 Multimodal complet / Full multimodal ✅
9 Vidéo / Video ✅
10 MoE multi-expert (20, top-2) ✅
11 Mémoire / RAG + apprentissage permanent ✅
12 API + Démo + Publication ✅

Détail complet : docs/phases.md. Choix techniques : docs/architecture.md. Journal des décisions : docs/decisions.md.


Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support