How to Use PSSM Generation
Commercially Available Online Web Server
Build and inspect reusable protein profiles from uploaded or generated MSAs.
Choose Upload an existing MSA for aligned FASTA/A3M, or Generate an MSA from sequences/query to align two or more supplied sequences with MAFFT or find homologs for one query with MMseqs2. Both modes use the same profile settings and results page.
How PSSM Generation Works
The profile stores background-smoothed probabilities and computes log2(probability/background) in bits. Default Henikoff weighting reduces redundancy; one total pseudocount is distributed over a uniform background per position. Custom background frequencies, row weights, ambiguity policy, query row and query-gap removal are available under Advanced Options. B/Z/J/X can be distributed; U/O must be ignored explicitly. Lowercase A3M residues and dots are insertion annotations excluded from match columns; lowercase FASTA residues are ordinary residues. Query coordinates retain original residue numbering, counting A3M insertions. CSV and viewer positions are one-based; reusable JSON coordinates are zero-based, with -1 for query gaps.
Download profile.json to retain exact probabilities, weighted counts, background, amino-acid order, sequence weights, coordinates and generation settings. Rounded CSV scores are for inspection and cannot reconstruct the profile reliably. Conservation is one minus observed weighted residue entropy divided by log2(20), before pseudocount smoothing. Coverage and gap fraction are unweighted across MSA rows. A position is flagged Low when fewer than three rows have a residue, weighted residue support is below two, or coverage is below 50%. These are heuristic flags, not confidence probabilities. Effective sequence count is (sum weights)² / sum(weights²). A single-sequence profile cannot provide evolutionary evidence. Scores are log-odds in bits, with no calibrated E-values or PSI-BLAST equivalence.
Reuse and downloads
Download profile.json and upload it to PSSM Sequence Alignment, PSSM Sequence Scan, or PSSM Profile Comparison. It retains full-precision probabilities, weighted counts, normalized background frequencies, sequence weights and their row names, original coordinates, the retained query and the generation settings. The version-1 JSON identifies format as neurosnap.pssm, version as 1, amino_acid_order as ACDEFGHIKLMNPQRSTVWY, and score_units as bits. Probabilities and counts have one row per retained position and 20 columns in that order; background has 20 entries; sequence_weights has one value per input MSA row. msa_columns and query_positions store zero-based original coordinates, with -1 for a query gap or no selected query. Scores can be reconstructed as log2(probabilities/background); zeros in probabilities yield negative-infinity scores.
The scores.csv, probabilities.csv and counts.csv tables are rounded to three decimals for inspection. Use profile.json for reuse or reconstruction. Negative-infinity scores appear as -inf in score tables and as blank cells in the heatmap; the profile JSON stores finite probabilities, including zeros. summary.csv reports row count, weight-based effective sequence count, retained length, low-support count and consensus. sequence_weights.csv pairs each row name with its weight; weights are normalized to sum to the MSA row count. Effective sequence count describes weight concentration, not independent evolutionary observations.
quality.csv reports position support, coverage, gaps and conservation. Non-gap counts include ambiguous characters even if the selected policy ignores their contribution; weighted residue support excludes gaps and ignored residues. All-gap or entirely ignored columns have conservation zero. The consensus is the most probable smoothed residue per position, with ties resolved by the amino-acid order; it does not establish function.
Download consensus.fasta for the profile consensus, source_alignment.a3m or source_alignment.fasta for the original alignment, and alignment.fasta for the rows restricted to retained profile columns. alignment_positions.csv maps every displayed residue to its original sequence, MSA and query positions with its residue score. The viewer previews up to the first 100 MSA rows, limited to 100,000 residue cells; download the complete table to inspect the remaining rows. Hover table headers or cells for column definitions.
Preparing the inputs
Fill the section for your selected Input Mode. Upload MSA accepts one FASTA or A3M file of up to 25 MB; generated alignments use supplied ungapped proteins or exactly one homolog-search query. The other input section is unused. The supported alignment size is up to 10,000 rows, 2,000 match columns and 2 million residue characters. When Query Row is 0, all columns are retained without query coordinates. Query-gap removal otherwise applies to the selected one-based row.
What is Neurosnap?
Neurosnap is the leading platform for bioinformatics and computational science focused on expanding access to powerful modeling and simulation tools. Because many state-of-the-art machine learning systems remain complex to install, configure, and scale, Neurosnap offers a clean, browser-based workspace that removes the burden of infrastructure management, dependency conflicts, and command-line tooling.
Built for biologists, chemists, and cross-disciplinary scientists, the platform enables advanced computational workflows without requiring expertise in software engineering or cloud architecture. Researchers can launch analyses through an intuitive interface, connect programmatically through a comprehensive API, and rely on automated resource management to scale workloads efficiently. By taking care of the underlying compute and operational complexity, Neurosnap allows teams to devote their energy to scientific progress and faster iteration. Security and data protection remain foundational principles, with clear safeguards outlined in our Terms of Use and Privacy Policy to ensure your work stays protected.
Advancing Discovery with PSSM Generation on Neurosnap
Using PSSM Generation on Neurosnap could drastically accelerate protein profile generation, inspection and candidate screening.
- Reusable profiles: Retain full-precision probabilities and source coordinates.
- Clear scoring: Interpret profile scores in bits.
- Downloadable results: Inspect tables and alignment details.
How to Use PSSM Generation on Neurosnap
To harness the capabilities of PSSM Generation, researchers can follow this streamlined workflow within Neurosnap:
- Access Neurosnap: Start by logging in to the Neurosnap website.
- Select Tool: From the list of available tools, choose PSSM Generation.
- Provide Inputs: Provide all the inputs specified within the submission panel and optionally configure the tool as desired.
- Run Tool: Submit the PSSM Generation job and Neurosnap will execute it in the cloud, automatically notifying you as soon as your results are ready.
- Review Output: Explore your results through rich visualizations, including figures, plots, and interactive views designed to help you analyze findings with clarity and confidence.
Citations
Please cite the original work when using PSSM Generation in publications or research outputs.
|
Amani, Keaun, and Danial Gharaie Amirabadi. 2024. Neurosnap SDK Package. Software. https://github.com/NeurosnapInc/neurosnap. |
|
Neurosnap Inc. (2022). Neurosnap: An online platform for computational biology and chemistry. Available at: https://neurosnap.ai/ |
Similar Services
Explore related tools that support similar research workflows:
Proudly supporting 50,000+ scientists worldwide, including 7,000+ leading biotech and global biopharma organizations.
Making Scientific Research
Faster & Easier
Register for free — upgrade anytime.
Interested in getting a license? Contact Sales.
Try Free