ainotis

What actually shipped. Every claim carries the source it rests on.

Research

Paper describes PrePssmCas, a classifier that fuses language-model and PSSM features for Cas proteins

2026-09-27 · that day's edition

Published as an accepted early version in BMC Bioinformatics, the method improves on the prior best tool by 3.91 percent in accuracy on an independent validation set.

Researchers described PrePssmCas, a machine-learning classifier for CRISPR-Cas protein systems that fuses two feature families: embeddings from pre-trained protein language models and position-specific scoring matrix (PSSM) features. The final classifier uses a 143-dimensional feature vector, 87 features from language-model embeddings and 56 from the PSSM representation. On an independent validation set the method reached 97.98 percent accuracy and a Matthews correlation coefficient of 0.962, improving on the previous best method, CRISPRCasStack, by 3.91 percent in accuracy and 9.60 percent in MCC. The paper appears as an accepted early version in BMC Bioinformatics.

Your reaction

One press adds one. We count a number for each notice and day, never who pressed it.

What this rests on

  1. On an independent validation set, PrePssmCas reached 97.98 percent accuracy and a Matthews correlation coefficient of 0.962.

    PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification · BMC Bioinformatics · 2026-09-26
    The selected features achieved an accuracy of 97.98% and an MCC of 0.962 on an independent validation set

  2. The authors report improvements of 3.91% in accuracy and 9.60% in MCC over CRISPRCasStack.

    PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification · BMC Bioinformatics · 2026-09-26
    representing improvements of 3.91% in accuracy and 9.60% in MCC over the previous best method, CRISPRCasStack.

  3. The final classifier uses a 143-dimensional feature vector: 87 features from pre-trained language-model embeddings and 56 from the RPM-PSSM representation.

    PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification · BMC Bioinformatics · 2026-09-26
    By employing random forest-based feature selection, a 143-dimensional feature vector was obtained, comprising 87 pre-trained features and 56 RPM-PSSM features.

We checked every sentence above against its source by opening it. Nothing appears on this site that we have not opened and linked.

Also that day