Paper describes PrePssmCas, a classifier that fuses language-model and PSSM features for Cas proteins
2026-09-27 · that day's edition
Published as an accepted early version in BMC Bioinformatics, the method improves on the prior best tool by 3.91 percent in accuracy on an independent validation set.
Researchers described PrePssmCas, a machine-learning classifier for CRISPR-Cas protein systems that fuses two feature families: embeddings from pre-trained protein language models and position-specific scoring matrix (PSSM) features. The final classifier uses a 143-dimensional feature vector, 87 features from language-model embeddings and 56 from the PSSM representation. On an independent validation set the method reached 97.98 percent accuracy and a Matthews correlation coefficient of 0.962, improving on the previous best method, CRISPRCasStack, by 3.91 percent in accuracy and 9.60 percent in MCC. The paper appears as an accepted early version in BMC Bioinformatics.