JCSG Technologies

JCSG's high-throughput Structural Genomics pipeline has delivered more than 1,500 protein structures to the community. Targets are processed through bioinformatics and biophysical analyses to characterize and optimize each one prior to structure determination, with parallel processing at almost every step. The pipeline adapts to a wide range of targets — from bacteria to human, including challenging eukaryotic proteins and macromolecular complexes — and has produced innovative methods, software, and free-access web-based tools (XtalPred, the Validation and Ligand servers, and the TOPSAN annotation portal).

From the outset, JCSG has been committed to developing new technologies and methodologies that advance high-throughput structural biology. Hardware development focuses on robotics that accelerate every stage of production; software development spans enterprise resource tracking, target management, and helper applications.

Target Selection

Genome Pool Strategy

Even closely homologous proteins often have different crystallization properties. The genome pool strategy targets multiple homologs in parallel, exploiting the broad distribution of physicochemical properties within most protein families to improve structural coverage — particularly for difficult families.

Publications

  • Jaroszewski L, Slabinski L, Wooley J, Deacon AM, Lesley SA, Wilson IA, Godzik A. Genome Pool Strategy for Structural Coverage of Protein Families. Structure 11, 1659–1667 (2008). PMID: 19000818.

XtalPred Server

Web server for predicting protein crystallizability. Compares protein features against TargetDB distributions, summarizes likely problems, predicts ligands, and (optionally) lists close homologs from microbial genomes more likely to crystallize.

Website: ffas.godziklab.org/XtalPred-cgi/xtal.pl

Publications

  • Slabinski L, Jaroszewski L, Rychlewski L, Wilson IA, Lesley SA, Godzik A. XtalPred: a web server for prediction of protein crystallizability. Bioinformatics 23, 3403–3405 (2007). PMID: 17921170.

Protein Production

Polymerase Incomplete Primer Extension (PIPE) Cloning

A rapid, efficient cloning process combined with protein microscreening (LC-MS, AnSEC) to evaluate suitability of constructs for X-ray crystallography. Used to clone 448 targets and generate 2,143 truncations from 96 targets, increasing structure determination success on recalcitrant targets by at least 38%.

Publications

  • Klock HE, Koesema EJ, Knuth MW, Lesley SA. Combining the polymerase incomplete primer extension method for cloning and mutagenesis with microscreening to accelerate structural genomics efforts. Proteins 71(2): 982–994 (2008). PMID: 18004753.

Microexpression System

High-throughput E. coli expression screening at small scale (~750 µL in 96-well blocks) reaching OD 10–20 with IMAC-based solubility evaluation. 97% of targets soluble in micro-expression scale up successfully to large-scale fermentation; adaptable to SeMet and 15N/13C labeling.

Cloning Robotics

Automated conventional cloning platform integrating liquid/plate handling, thermocyclers, and a plate reader — capable of producing up to 384 validated expression clones per week from a single operator. Over 2,500 expression clones generated to date.

Large-scale Bacterial Expression (GNFermentor)

Parallel 96-culture high-density fermentation system producing 2–4 g of cell pellet per culture with <5% pre-induction OD variation. Tightly regulated arabinose induction; over 30,000 individual samples processed.

Publications

  • Kreusch A, Lesley SA. High-Throughput Cloning, Expression, and Purification Technologies. Genomics, Proteomics, and Vaccines, ed. G. Grandi, Wiley Press, 171–184 (2004).

Automated Mammalian and Insect Cell Expression (PEPP)

The Protein Expression and Purification Platform (PEPP), developed at GNF, provides a robust HT vehicle for evaluating eukaryotic-protein expression constructs in mammalian and baculovirus systems for crystallization.

Automated Affinity Purification (GNFuge)

Direct processing of fermentation tubes for lysis, debris removal, and affinity purification, feeding directly into secondary purification or crystallization screening.

Publications

  • Lesley SA. High-throughput proteomics: protein expression and purification in the postgenomic world. Protein Expr. Purif. 22(2): 159–164 (2001). PMID: 11437590.

Secondary Purification

Configured Akta Purifyer systems (GE Healthcare) with custom valves and air sensors enable automated processing of up to 12 samples without commercial-autosampler volume limits — ~48–96 proteins per week at 10–50 mg scale across three online systems.

Biophysical Characterization

Biophysical characterization is critical for guiding target strategy and evaluating pipeline performance. Parameters tracked per target:

Toxicity during expression

Final optical density.

Cofactor binding

UV/Vis absorbance scan.

Protein concentration

Bradford assay.

Protein purity

SDS-PAGE.

Isoelectric point

IEF gel electrophoresis.

Protein fingerprinting

Tryptic mass spectrometry.

Thermostability

Differential Scanning Calorimetry.

Polydispersity / Native MW

Analytic Size Exclusion Chromatography (AnSEC).

Metal binding

X-ray Absorption Fine-Structure (XAFS) spectroscopy.

Crystallization

Nano-drop Crystallization & Robotics

JCSG members pioneered nanoliter-volume crystallization using custom robotics from Syrrx/GNF — the first center to apply this technology worldwide. Imaging across two custom platforms in 4 °C and 20 °C rooms with 1,536-plate capacity has produced over 3,000,000 images on a 7/14/28-day schedule.

Publications

  • Santarsiero BD et al. An approach to rapid protein crystallization using nanodroplets. J. Appl. Crystallogr. 35, 278–281 (2002).

Rigaku CrystalMation Platform

The largest fully integrated HT crystallization platform in the U.S. — covering screen making, automated trials, imaging, and analysis. Sets up 100 96-well plates every 8 hours with a total 4,000-plate/month capacity. >9,000,000 images screened, >124,000 crystals harvested.

Crystal Screening for Diffraction Quality

JCSG produces over 500 crystals per month for diffraction screening. Automated screening, co-developed with the SSRL Structural Molecular Biology group, eliminates manual mounting bottlenecks.

Compact Crystal Cassette

Cylindrical aluminum cassette holds 96 crystals on Hampton Research pins. Two cassettes fit in a standard vapor shipping dewar; twenty fit in a Taylor-Wharton HC-35 storage dewar. Distributed widely to SSRL users.

Website: smb.slac.stanford.edu cassette kit

Publications

  • Cohen AE, Ellis PJ, Miller MD, Deacon AM, Phizacherley RP. An automated system to mount cryo-cooled protein crystals on a synchrotron beam line, using compact sample cassettes and a small-scale robot. J. Appl. Crystallogr. 35, 720–726 (2002).

Stanford Auto-Mounter (SAM)

Epson ES553S 4-axis robot with pneumatic cryo-tong mounts crystals from cassettes onto the goniometer; supports cassette-to-cassette sorting. Integrated with BLU-ICE; by end of 2008, >85% of SSRL PX experimenters used SAM.

Website: smb.slac.stanford.edu SAM

Sample Visualization & Loop Alignment

Navitar 12× lens with Optronics CCD on Axis 2400 web image server, plus diffuse backlight, enables ~30-second automated alignment with >95% reliability. A 96-crystal cassette can be screened unattended in ~5 hours.

Diffraction Data Collection

Automated MAD Data Collection with BLU-ICE

BLU-ICE supports completely automated MAD data collection — wavelengths derived from Kramers-Kronig analysis of fluorescence scans, automatic energy changes, intensity optimization, and dose-mode normalization across wavelengths.

Website: smb.slac.stanford.edu BLU-ICE

Publications

  • McPhillips TM et al. Blu-Ice and the Distributed Control System: software for data acquisition and instrument control at macromolecular crystallography beamlines. PMID: 12409628.

Remote Data Collection

Combined with SAM, the entire diffraction experiment can be initiated remotely. Live BLU-ICE video feeds aid remote diagnosis. Over 75% of SSRL PX experiments are now conducted remotely.

Publications

  • Soltis SM et al. New paradigm for macromolecular crystallography experiments at SSRL: automated crystal screening and remote data collection. Acta Cryst. D 64, 1210–1221 (2008). PMID: 19018097.

Data Processing & Structure Determination

Xsolve

Linux-based parallel processing environment that executes all crystallographic data processing and MAD structure determination steps — indexing, integration, scaling, phasing, density modification, initial model building — and prepares files for upload to the Structure Solution Tracking System (SSTS).

Customized Scripts

In-house scripts to prototype new programs and enable rapid data processing at remote synchrotron sources, including XDS+Solve and SHELX+Solve interfaces, made available to SSRL users.

Molecular Replacement (MR) Pipeline

Highly parallelized MR pipeline using FFAS03 profile-profile alignments and fold-recognition models. Successfully applied to 26+ cases <35% sequence identity, 10 cases <30%, and several near 15%.

Publications

  • Schwarzenbacher R, Godzik A, Grzechnik SK, Jaroszewski L. The importance of alignment accuracy for molecular replacement. Acta Cryst. D 60, 1229–1236 (2004). PMID: 15213384.
  • Schwarzenbacher R, Godzik A, Jaroszewski L. The JCSG MR pipeline: optimized alignments, multiple models and parallel searches. Acta Cryst. D 64(Pt 1), 133–140 (2008). PMID: 18094477.

Model Building, Refinement & Quality Control

Refinement Standards

All JCSG refinement uses the latest Refmac with TLS, riding hydrogen, and NCS restraints evaluated for R-free impact. Experimental phase restraints are always included when available. WHAT_CHECK, ADIT, PDB tools, and MolProbity validate every structure.

Validation Suite (QC Server)

JCSG Quality Control Server processes coordinates and data through AutoDepInputTool (Yang et al., 2004), MolProbity (Davis et al., 2004), WHATIF 5.0 (Vriend, 1990), RESOLVE (Terwilliger, 2003), MOLEMAN2 (Kleywegt, 2000), and several in-house scripts, summarizing results.

Website: smb.slac.stanford.edu/jcsg

Structure Deposition in the Protein Data Bank

Automated mmCIF Deposition

Coordinates passing QC are combined with database-derived target history and methods, parsed into two mmCIF files (release coordinates; structure factors and unmerged intensities with experimental phases), and deposited directly with the PDB. Largely automated, with manual oversight for completeness checks.

Computational Target Analysis & Functional Annotation

Unified Sequence/Structure Analysis

Includes structure similarity (DALI, CE, FATCAT), distant homology (FFAS), and pathway/genome context (SEED). Annotations curated through interactive pages — establishing functional annotations, including for previously unannotated hypothetical proteins, for over half of JCSG-solved targets.

Protein Sequence Comparative Analysis System (PSCA)

Tabbed interface to public-database annotations, links, and preprocessed target information including fold and sequence similarity, domain organization, and physicochemical properties — precalculated for fast access.

TOPSAN — Wiki-based Collaborative Annotation

Combines automated and expert-curated annotations from JCSG personnel and the wider community. Built on a wiki platform but oriented toward generating new knowledge through instant collaboration across distributed PSI participants.

Website: www.topsan.org

Publication & Data Dissemination

Public Tracking System & Website

Provides high-level target tracking, production metrics, bioinformatics, and visualization tools — including a graphical view of every target's full pipeline history. Serves as primary public outreach channel.

Customized Tracking Lists

Public tracking interface and weekly XML exports to TargetDB are auto-generated. Registered users can subscribe to email alerts on individual targets and create personalized JCSG database views.

Structure Notes

Short papers describing annotation, biology, structure, and functional implications of each protein, auto-assembled from the JCSG database (annotation, cloning, purification, crystallization, data collection, refinement statistics) with manual structural figures via PyMOL.

Downloadable Datasets

Repository of X-ray crystallographic datasets — data collection, reduction, phasing, density modification, model building, and refinement — available as test data to the methods development community.

Database & LIMS Development

Tracking Database

Built from scratch in Oracle: 140+ tables describing 32 production stages and tracking 530+ parameters. Perl interface with 1,800 custom scripts, 100 user interfaces, 30 daily reports in XML and Excel — about 360,000 lines of code.

Publications

  • Godzik A, Canaves J, Grzechnik S, Jaroszewski L, Morse A, Ouyang J, Wang X, West B, Wooley J. Challenges of structural genomics: bioinformatics. Biosilico 1, 36–41 (2003).

Laboratory Information Management System (LIMS)

Tracks every step from target activation to deposition, with submenus tailored to each core. Functions as the central hub directing information flow within JCSG.

Pipeline Data-Flow Analysis & Data Mining

PCR Amplification Analysis

Feedback from success-rate analysis improved the primer generation system, adding a scoring function for optimal GC clamps. Optimized primer sets achieve ~98% success rates.

Crystallization Screen Analysis (GNF96)

Analysis of 340,000+ trials yielded GNF96 — a 96-condition coarse screen capturing 84% of crystallizing proteins. Available commercially via Qiagen as JCSG+ Suite and the JCSGCore Suites.

Publications

  • Lesley SA, Wilson IA. Protein production and crystallization at the Joint Center for Structural Genomics. J. Struct. Funct. Genomics 6(2-3), 71–79 (2005). PMID: 16211502.