PROJECT 14 / 18SCIENTIFIC VISUALIZATIONPYTHON

Scientific visualization toolkit

Protein Viz.

Turn a structure file into something you can explore.

6VYBdefault structure ID
200default Cα residue cap
4generated molecular view modes
01 / IDEA02 / SYSTEM03 / PLAYGROUND04 / DECISIONS05 / SOURCE
01 / THE IDEA

A closer look.

A Python pipeline for retrieving protein structure files, analyzing their coordinates, and generating static plots plus interactive molecular viewers. Its default example uses PDB 6VYB, with BioPython parsing and a vectorized Cα distance matrix.

A coordinate file contains far more information than its rows reveal at a glance. Linked views of sequence composition, spatial distances, annotated regions, and deposited atom values make that structure easier to inspect.

01

Coordinate analysis

The analyzer computes chain statistics, residue composition, biochemical property groups, and per-residue B-factor means from the parsed structure.

02

Vectorized geometry

The contact-map kernel collects Cα coordinates from one chain, then uses NumPy broadcasting to compute all pairwise Euclidean distances.

03

Multiple views of one structure

The pipeline writes static charts and separate HTML molecular viewers for cartoon, surface, region, and deposited B-factor views.

04

Inspectable output

Source coordinate files are cached locally and the analysis summary is exported to JSON, making the plotted values easier to trace back to their inputs.

02 / UNDER THE SURFACE

Coordinates become views

Follow the downloaded structure through model selection, analysis kernels, and independent output artifacts.

DRAG TO PAN · SELECT A NODE · + / − TO ZOOM

Read the architecture as text
  1. Structure selection — The CLI accepts a PDB identifier and a chain for the distance matrix. The default is 6VYB, chain A. Other analyses in run_pipeline operate across the first model’s chains.
  2. Structure cache — fetch_pdb_file normalizes the ID, returns an existing local file unless forced, or downloads the requested PDB/mmCIF file from RCSB.
  3. Deposited metadata — A separate REST request extracts deposited metadata including experimental method, resolution, counts, and citation information. These values are distinct from counts computed from parsed atoms.
  4. BioPython structure — parse_structure reads the cached structure into a BioPython hierarchy. The analysis functions explicitly select structure[0], the first parsed model.
  5. Standard amino acids — Composition and residue-profile calculations include standard amino acids. Chain atom counts still count all atoms, so glycan or other non-protein chains may have atoms but no amino-acid residues.
  6. Residue B-factors — get_residue_bfactors averages atom B-factor fields within each standard amino-acid residue and keeps chain/residue identifiers. These are deposited structural values, not a molecular dynamics simulation.
  7. Composition + properties — Standard residue names map to one-letter codes and counts. A deterministic category chain assigns each residue to hydrophobic, polar, positive, negative, or special groups.
  8. Cα coordinate selection — compute_ca_distance_matrix selects standard residues containing a CA atom in the chosen chain. It stops after max_residues, which defaults to 200.
  9. Pairwise distances — NumPy broadcasting forms all pairwise coordinate differences. Squared components are summed and square-rooted to produce a symmetric Euclidean distance matrix in coordinate units.
  10. Known region annotations — identify_key_regions returns curated residue ranges keyed by the structure ID. The current table includes 6VYB; unknown IDs return an empty annotation list.
  11. Structural summary — summarize_structure combines chain statistics, composition, property distribution, and known regions. The pipeline writes that summary as an analysis JSON file.
  12. Static plots — The pipeline generates residue profiles, composition and property charts, chain comparisons, a domain map when annotations exist, and a distance heatmap when its calculation succeeds.
  13. Browser molecular scene — The viewer generator writes HTML with 3Dmol.js, fetches the PDB data in the browser, and configures cartoon, VDW surface, region, or B-factor coloring.
  14. Four view modes — create_all_views emits four HTML pages. Region highlighting uses configured spike ranges on chain A; generated viewer titles and annotations are specialized to the spike example.
03 / INTERACTIVE STUDY

From a chain to a contact map

Explore an illustrative Cα backbone and its pairwise distances. Adjust the number of residues, a display contact cutoff, and the viewing angle.

CHANGE THE INPUTS

Synthetic coordinates only. The source computes a full distance matrix; the cutoff is an explanatory display filter. Rotating the view must not change pairwise distances.

ILLUSTRATIVE MODELLIVE

04 / ENGINEERING CHOICES

Why it works this way.

01

Bound the contact matrix

The default cap limits the selected chain to its first 200 Cα-bearing standard residues. This keeps the quadratic distance matrix small, while making its scope explicit.

02

Separate deposited and computed counts

RCSB metadata and parsed coordinate counts answer different questions. The pipeline obtains deposited metadata independently from its own standard-residue analysis.

03

Annotate known structures explicitly

The source uses a static region table for 6VYB, rather than inferring biological function from coordinates. Additional PDB IDs do not automatically gain validated annotations.

05 / OPEN THE SOURCE

Trace it back.

Implementation details, examples, and project documentation.

Scope & limitations

  • The portfolio interaction uses a synthetic coordinate example; it is not the actual 6VYB structure, protein folding, molecular dynamics, or a medical tool.
  • Analyses select the first parsed model. The contact map has a default residue cap, and viewer region annotations are specialized to the spike example.

Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.