JCSG Tool

Protein Sequence Comparative Analysis (PSCA)

PSCA is an integrated web tool for comparative analysis of protein sequences. It analyzes sequences in multiple layers — domains and families, secondary structure features, similarity to PDB protein structures, structure prediction, and homologue search — using the JCSG server and database infrastructure.

Historical note. PSCA was built and operated during the JCSG program (2000–2015). This page documents the tool's methodology and the external databases it drew on; the original interactive JCSG server may no longer be available. The third-party resources it integrated (NCBI BLAST, Pfam, InterPro, RCSB PDB, and others) remain publicly accessible at the links below.

About PSCA

Protein Sequence Comparative Analysis (PSCA) is an integrated web tool for comparative analysis of protein sequences. It analyzes a protein sequence across multiple layers — domains and families, secondary structure features, similarity to PDB protein structures, protein structure prediction, and protein homologues. The PSCA user interface is supported by the server and database of the Joint Center for Structural Genomics (JCSG).

Layers of Comparative Analysis

Sequence Information

Protein sequence information aggregates annotation content from JCSG and SWISS-PROT. SWISS-PROT and TrEMBL annotations are accessed via the EBI SRS server (SWALL). Enzymatic and metabolic information for enzyme targets is drawn from KEGG. JCSG-internal tools — including the Data Acquisition Prioritization System (DASP), Functional & Structural Space (FSS), and Target PDB Monitor (TPM) — provide additional target annotation and coverage tracking.

Homologous Protein Sequences

PSCA uses NCBI BLAST and PSI-BLAST to search homologues against multiple sequence sets with an E-value threshold of 0.001. Search sets include NCBI NR and JCSG sets such as JCSG Target, Thermotoga, Yeast, C. elegans, and Mouse.

Domains and Families

Domain search uses the HMMER program against the Pfam database. PSCA renders Pfam domain hits in both graphic and textual formats — single domains in a single color across the matched subsequence, overlapping domains as mixed colors. Users can also retrieve a specified Pfam domain across a JCSG subset such as the Thermotoga maritima or Caenorhabditis elegans proteomes.

Secondary Structure Feature

JNET predicts secondary structure elements (alpha helices and beta strands). TMHMM predicts transmembrane helices using a hidden Markov model.

PDB Fold Similarity

Homologous sequences typically share similar structures. PSCA uses BLAST to search the non-redundant PDB sequence database (pdbnr), built from representative PDB chains chosen for the best percent of structural coverage (%covp). Because PDB structures may be partially solved or contain missing residues, %covp reflects the difference between the submitted sequence and the atom sequence extracted from the coordinate file: covp (%) = (aligned atom residues − gaps) / (length of real sequence) × 100%.

Fold & Function Assignment System (FFAS)

FFAS is a profile–profile fold recognition method developed by the Godzik Laboratory, used by PSCA to extend recognition of distant structural homologues beyond standard sequence alignment.

Reference System

PSCA collects literature references in two ways: a subject-oriented search using protein description and keywords, and an automated sequence-oriented search driven by protein sequence annotation. References are gathered from the target sequence itself, its domains and families, PDB structures, and NCBI NR homologues. Selection controls let users filter references by interest or by automatically computed relevance. Entrez PubMed provides access to MEDLINE citations.

Reference Classification

References are classified into three groups using transparent criteria. The query sequence itself is treated as Trusted. Pfam domains and families are scored against Pfam HMM thresholds (Trusted / Gathering / Noise). InterPro domains and families collected from EBI SWALL (SPTR) and EBI InterPro are treated as Trusted. BLAST homologues are scored as Extreme similarity (identity ≥ 85% and coverage ≥ 50%), High similarity (identity ≥ 30% and coverage ≥ 50%), or Low similarity (E-value ≤ 0.001 but below high-similarity thresholds). For PSI-BLAST, the high-similarity identity threshold is relaxed to 25% because PSI-BLAST alignments are typically longer and more sensitive. By default the reference list is shown at the Gathering / High similarity level; users can rebuild the list using filter controls.

Genome Reference Filter

Some genome-scale references are non-specific for a given target. The genome reference filter removes these non-specific genome references from the displayed list.

How To Use PSCA

Direct PSCA server call

Given a protein ACCESSION, you can analyze the sequence with or without an explicit threshold (E-value cut-off). When no threshold is provided, PSCA uses default expects of 0.01 for Pfam domain search and 0.001 for BLAST homologue search. Set hexpect for Pfam and bexpect for BLAST to override.

Acknowledgement

Homologous Sequences

Altschul SF, Madden TL, Schäffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997;25(17):3389–3402. Sequence clustering uses CD-HI (Li W, Jaroszewski L, Godzik A. Clustering of highly homologous sequences to reduce the size of large protein databases. Bioinformatics. 2001;17:282–283), developed by the Godzik Laboratory.

Pfam Domains

Bateman A, Birney E, Durbin R, Eddy SR, Howe KL, Sonnhammer ELL. The Pfam protein families database. Nucleic Acids Res. 2000;28:263–266. Domain and family search performed with HMMER against Pfam at an E-value cut-off of 0.01.

Secondary Structure Feature

Cuff JA, Barton GJ. Application of enhanced multiple sequence alignment profiles to improve protein secondary structure prediction. Proteins. 1999;40:502–511 (JNET). Transmembrane helix prediction uses TMHMM, a hidden-Markov-model method developed by Anders Krogh and Erik Sonnhammer.

PDB Fold Similarity

PDB sequences and structure data are downloaded from the RCSB PDB FTP server. BLAST searches run against the pdbnr non-redundant PDB sequence database, built from representative chains with the best structural coverage in each sequence-identity group.

Fold & Function Assignment System (FFAS)

Jaroszewski L, Rychlewski L, Godzik A. Improving the quality of twilight-zone alignments. Protein Sci. 2000;9(8):1487–1496. FFAS is developed and maintained by the Godzik Laboratory.

Reference System

Entrez PubMed provides access to MEDLINE citations. Sequence information is collected from NCBI GenBank and SWALL (SPTR) on the EBI SRS server. Functional and structural information is collected from EBI InterPro, Pfam, and the RCSB PDB.