Hugging Face Releases Carbon-A, a 1.2B Gene Finder, and 566 Million Gene Candidates Across 22,617 Species
Summary
Hugging Face's HuggingFaceBio team releases Carbon-A, a 1.2-billion-parameter model that predicts protein-coding genes from DNA, and the Carbon Annotation Database of 566 million gene candidates across 22,617 species. The team reports a 0.944 nucleotide F1 across 42 benchmark genomes. Iso-Seq tests in cat, hamster, chicken and Arabidopsis support some predictions missing from RefSeq.
Key Points
- Carbon-A processes a 98,304-base-pair context window and makes nucleotide-resolution predictions on both DNA strands, according to Hugging Face.
- Fergal Martin of EMBL-EBI says the dataset could help researchers interpret and compare genes and genomes at scale, and EMBL will soon host the annotations.
- Hugging Face plans to release another batch of GenBank annotations in three weeks, after annotating roughly half of its target genomes.