RCSB protein data Bank: Next-generation advances search for exploration of experimental structures and computed structure models

Author(s)

Rose Y, Brown R, Voigt M, Bhikadiya C, Bittrich S, Duarte JM, Segura J, Green RK, Burley SK. RCSB protein data Bank: Next-generation advanced search for exploration of experimental structures and computed structure models.

Sources

Protein Sci. 2026 Aug;35(8):e70731. doi: 10.1002/pro.70731. PMID: 42517697; PMCID: PMC13410952.

The Protein Data Bank (PDB), established in 1971, is the primary global, open-access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research-focused RCSB.org webportal provides access to these data alongside more than one million machine-learning-predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community.

Unified Query Builder: Seamlessly combines annotation filters, sequence similarity search and 3D shape similarity into complex Boolean queries.

Integrated 3D Viewer: Features an embedded Mol* 3D tool inside the query builder, allowing you to define structural spatial parameters without needing formal nomenclature or residue numbers

Curated Biological Knowledge: Builds geometry-driven queries using active-site templates from the Mechanism and Catalytic Site Atlas (M-CSA) and ligand-guided structural motifs.

Advanced Chemical Search: Features an integrated chemical drawing tool to query small molecules via standard identifiers (SMILES, InChI, Chemical Formula), linked directly to structural annotations

Cross-Model Exploration: Allows users to query purely experimental data, include AI/ML models (AlphaFold DB and ModelArchive), or isolate Integrative/Hybrid Methods (IHM) structures.

The update significantly improves how results are refined, organized, and retrieved:

  • Result Clustering: Results can be automatically grouped into UniProt Groups, Sequence Identity clusters, or Deposition Groups to map trends across related structures.
  • MyPDB Integration: Users can authenticate via Google, Facebook, or ORCID to store search histories, re-run complex workflows, and subscribe to automatic email updates for weekly data releases.
  • Programmatic Support: Advanced graphical queries built in the user interface can be exported into REST and GraphQL API queries via a dedicated interface

Latest news

This study uncovers a surprising, non-linear behavior in how polysaccharides change shape when you add...

Isotopic metabolomics reveals that plant species exhibit highly divergent carbon allocation and metabolic rewiring strategies...

Carbohydrates are among the most abundant and structurally diverse biomolecules in nature, playing central roles...

Pushing Frontiers for Proteoglycans” is a landmark scientific perspective article published in 2026 that details...