Carbon-A: Finding genes in known and unknown genomes
You can explore the Carbon Annotation Database, the model, the training data, and the technical report in HuggingFaceBio’s Carbon Annotation Database Collection.
Why working on models for biology now?
We’re seeing signs that the major AI labs are gearing up to tackle biology as the next frontier after cyber and math. We think 2027 will be the year for bio. It has the potential to improve our lives far beyond what we’ve seen from AI in other fields so far.
But we’ve also seen how fragile the ecosystem can become when researchers depend on access to closed frontier models. The controversy around Navier-Stokes and Claude refusing legitimate biology questions raise concerns about who controls these tools and the research they enable.
To keep the open science ecosystem going, we want to help make these models and the insights they produce accessible, so scientists can use, study, and build on them without depending entirely on the priorities and policies of a handful of closed labs.
Introduction
Sequencing a genome means physically reading out an organism’s DNA—the billions of A, C, G, and T bases that make up its chromosomes. For decades, DNA had to be broken into manageable pieces, read through relatively slow and expensive laboratory processes, and then computationally reassembled into a genome. Producing a high-quality genome could take years and cost enormous amounts of money.
Fortunately, that has changed dramatically. Modern sequencing technologies can read huge amounts of DNA in parallel, while improvements in chemistry, instrumentation, and computation have driven the cost and time of sequencing down by orders of magnitude. We can now generate genome assemblies far faster than we can understand them.
The bottleneck has therefore moved downstream to genome annotation: figuring out where the genes are, which regions encode proteins, and how those coding regions are organized into exons and introns. We focus on protein-coding regions, which are stretches of DNA that contain the biological instructions for making proteins. Proteins shape how cells work and can reveal traits, adaptations, and biological pathways across species. Finding genes is an early step in turning raw genome sequence into biological insight, from understanding evolution and ecology to discovering useful molecules, enzymes, and potential drug targets. However, annotation still usually depends on transcript data, known proteins, interpolation from closely related species, and expert involvement, which, to top it off, is all unevenly distributed across the tree of life.
Now, we’re releasing Carbon-Annotator (Carbon-A), a 1.2B-parameter model that predicts protein-coding regions directly from DNA using a single model across mammals, other vertebrates, invertebrates, plants, fungi, and protists. It works over a 98,304-base-pair context window and produces nucleotide-resolution predictions on both strands.
Across 42 benchmark genomes, Carbon-A reaches a macro-averaged nucleotide F1 of 0.944 and outperforms the baselines we evaluated (AUGUSTUS, Helixer, Tiberius, ANNEVO, OrionGeno, SegmentNT, and NTv3) at nucleotide, exon, and gene level.
Learning from imperfect labels
Carbon-A was trained on RefSeq, NCBI’s bank of reference genome annotations. However, though we use these labels as our ground truth in training, these reference annotations are not static! Genome sequences become more accurate, new experimental evidence accumulates, and annotation errors are corrected over time. This creates an interesting challenge because the model has to become robust to errors in its training labels.
Let's consider the pig genome.
Our training snapshot included an older pig assembly from 2017. After training, a newer high-quality pig assembly was released in 2026. Carbon-A performs better against the new assembly than the older one. That suggests the model is not simply memorizing its training annotations. Instead, it appears to have learned coding regularities that remain useful as assemblies and annotations improve.
One model across the tree of life
Most eukaryotic gene finders are specialized by species or clade. Carbon-A instead uses one shared model across major eukaryotic groups. To test how far that representation transfers, we evaluated an animal-only Carbon-A checkpoint on plants it had never seen during training. It still does a remarkable job annotating the plant genomes! For more detail on this, please check out our technical report.
We also tested Tetrahymena thermophila on an out-of-distribution case, which uses a nonstandard genetic code where two usual stop codons encode glutamine. Carbon-A still maintains extremely high accuracy (0.960 nucleotide F1) without requiring a species-specific annotation model. These stress tests suggest that Carbon-A is learning transferable features of coding sequence rather than only lineage-specific heuristics.
Wet Lab Validation
The interesting results are the ones where the model disagrees with the reference database: are detected genes novel discoveries or just false positives? We ran some initial wet lab results to verify newly predicted genes.
To get an independent view, we worked with researchers at ActiveSite and UCSD to run PacBio Iso-Seq, a lab technique that reads full, real-life RNA molecules straight from actual living cells. If Iso-Seq detects a gene structure, it’s strong independent evidence that the transcript structure is real.
We compared Carbon-A, RefSeq, and several other gene predictors against Iso-Seq data from cat, Syrian hamster, chicken, and Arabidopsis. The Iso-Seq data support several coding regions predicted by Carbon-A that are absent from the RefSeq annotation, providing independent evidence that some apparent disagreements with the reference may reflect missing annotations rather than model errors. In the aggregate results below, Carbon-A reaches roughly 0.62 complete-CDS support, which closely matches the reference closely and outperforms all other ab initio systems we tested.
For us, this is an important complement to reference-based benchmarks. Reference annotations and model predictions are both hypotheses about underlying biology, but the independent experimental measurements can help understand and resolve the disagreements.
These experiments provide evidence that the predicted regions are transcribed into RNA, and support their exon structure. They do not yet establish that those RNAs are translated into proteins—or what those proteins do. The next step is to look for evidence of translation using ribosome profiling and to detect the predicted proteins through mass spectrometry. Targeted experiments can then test their function, moving from a candidate gene to an understanding of its role in the organism.
We have additional wet-lab results that we can’t share yet, but more on those soon!
Annotating GenBank at scale
To make this model more useful for the genomics community, we have also been running Carbon-A across eukaryotic assemblies in GenBank to build the Carbon Annotation Database (CADB).
The current release contains 48,167 assemblies from 22,617 taxa, covering around 27 trillion base pairs and 566 million predicted protein-coding loci. Compared with the RefSeq-derived training corpus, that is about 11.0× more taxa and 9.0× more sequence.
Each prediction is linked to its assembly, genomic coordinates, reconstructed coding sequence, translated protein, and a confidence score, so researchers can filter the data according to their needs.
We have now annotated roughly half of the genomes in our target GenBank set, and we plan to release another batch in three weeks.
These annotations will also soon be available through EMBL, and they, like us, are excited about the impact of this kind of ML model for advancing genomic understanding!
This new dataset was built on the huge number of publicly available genomes in the European Nucleotide Archive and could help researchers interpret and compare genes and genomes at scale,” said Fergal Martin, Eukaryotic Annotation Team Leader at EMBL’s European Bioinformatics Institute (EMBL-EBI). “Such work highlights the different opportunities for applying AI to genome annotation.
Run it yourself
If your species is not in the database yet, you can run Carbon-A directly from the Hugging Face Hub: https://huggingface.co/HuggingFaceBio/Carbon-A-1.2B
If you want to explore the genomes we've already annotated, check out the CADB Explorer: https://huggingface.co/spaces/HuggingFaceBio/genbank-annotation-explorer
As a bonus: The released model also provides a confidence score for each predicted gene. On our benchmark, this score reaches an AUROC of 0.876 for distinguishing exact CDS matches from other predictions. This makes it possible to trade coverage for precision depending on the application.
What comes next
Our next steps include:
- gathering feedback from researchers working on poorly annotated genomes;
- exploring models that predict complete gene structures more directly;
- extending Carbon-A to handle multiple isoforms explicitly;
- continuing experimental validation, including proteomics and mass spectrometry;
- and continuing to expand the Carbon Annotation Database.
The current model already shows signals associated with alternative isoforms, but the decoder still produces a single consensus path per locus. Resolving multiple transcripts explicitly is one of the clearest next steps.
The broader goal is simple: public databases already contain an enormous amount of genomic sequence from poorly studied organisms. Carbon-A is an attempt to make more of that sequence interpretable—and turn it into useful hypotheses, mechanisms, and ultimately interventions over genes and proteins.


