Free during beta

Will your protein survive native MS?

Paste a sequence. Get a calibrated suitability score in seconds. Save days of failed experiments.

Predict native MS suitability

Paste a single protein sequence below. FASTA headers are fine — we'll strip them.

Try with example:

50 free predictions / month / IP. No login required.

How it works

Three steps, no setup. Built for bench scientists who want a sanity check before spending a week on sample prep.

01 INGEST

Paste a sequence

Drop in a single protein in plain text or FASTA. We clean and validate it client-side before sending anything.

02 ANALYZE

Score it with a foundation model

The v0.4 model combines 47 hand-engineered features (physicochemical + N-glycosylation sequon counts + transmembrane-helix prediction) with mean-pooled ESM-2 protein-language embeddings, trained on 635 proteins from PDB and EuropePMC.

03 DELIVER

Get a score and clear next steps

You see a suitability score from 0 to 100, the specific risk factors that hurt it, and concrete buffer and instrument recommendations.

About

Open benchmark and triage model for native MS suitability. v0.4 trained on 635 proteins from public datasets (PDB, UniProt, EuropePMC), with explicit N-glycosylation sequon and transmembrane-helix features. The risk panel surfaces membrane-topology and glycosylation signals with targeted experimental recommendations (nanodiscs / amphipols / SMA for membrane proteins; PNGase F deglycosylation for glycoproteins). Code, dataset, and trained model on GitHub. Methodology and benchmarks: bioRxiv preprint (DOI 10.64898/2026.05.03.722506).

What this is. A first-pass triage tool. Useful for ranking candidates before committing instrument time. Not a substitute for empirical optimization, and not a guarantee of success.
Known limitations. Trained on 635 proteins (538 positives plus 97 negatives, of which 3 are evidence-based real failures and 94 are proxy / property-targeted records). v0.4.1 stratified cross-validated AUC = 0.870 +/- 0.037, cluster-aware AUC = 0.835 +/- 0.029. Failure-detection performance is not yet statistically validated. Sequences over 1,022 amino acids are truncated for the ESM-2 component. The model is most reliable as a positive-suitability triage tool: high-confidence positive predictions can be trusted; low-confidence predictions should be treated as a flag for manual review, not a verdict. The membrane-topology and glycosylation risk rows are model-aware signals to consider alongside the score, especially when the score itself looks high but topology indicates special handling is needed. Treat scores as guidance, not as definitive predictions.