MCP server to search, inspect, and cross-reference RCSB Protein Data Bank structures
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
An MCP server for interrogating Protein Data Bank structures β discover, inspect, and cross-reference β from LLM clients (Claude Desktop, MCP Inspector, Cursor, etc.). It spans three RCSB APIs:
| Tool | What it does |
|---|---|
rcsb_list_pdb_search_attributes | Discover searchable attribute paths, types, and operators. schema="structure" (default, ~677) or schema="chemical" (~57: chem_comp.*, drugbank_info.*, ...). |
rcsb_find_go_terms | Resolve a free-text molecular function / biological process / cellular component to Gene Ontology ids (via EBI QuickGO), annotated with PDB entry counts β then search by rcsb_polymer_entity_annotation.annotation_lineage.id. |
rcsb_find_interpro_domains | Resolve a free-text protein domain / family / fold to InterPro ids (via EBI InterPro API), annotated with PDB entry counts β then search by rcsb_polymer_entity_annotation.annotation_id. |
rcsb_find_enzyme_classes | Resolve a free-text enzyme / reaction to Enzyme Commission (EC) numbers (via EBI Search/IntEnz), annotated with PDB entry counts β then search by rcsb_polymer_entity.rcsb_ec_lineage.id (hierarchical). |
rcsb_find_disease_terms | Resolve a free-text disease / condition to MONDO ids (via EBI OLS), annotated with PDB entry counts β then search by rcsb_uniprot_annotation.annotation_lineage.id (hierarchical, UniProt-based). |
rcsb_find_organisms | Resolve a free-text organism / common name / clade to NCBI Taxonomy ids (via UniProt taxonomy), annotated with PDB entry counts β then search by rcsb_entity_source_organism.taxonomy_lineage.id (hierarchical: a clade id matches every organism beneath it). |
rcsb_search_fulltext | Free-text keyword search (e.g. "CRISPR Cas9"), optionally refined with structured attributes filters (AND/OR) and sort. |
rcsb_search_by_attribute | Structured search on one or more indexed attributes (resolution, organism, release date, ...) combined with a single AND/OR. Each AttributeFilter supports exists, negation, case_sensitive; chemical=True (text_chem). |
rcsb_search_by_sequence | MMseqs2 sequence-similarity search (BLAST-like). |
rcsb_search_by_chemical | Chemical search by SMILES/InChI descriptor (whole-molecule or substructure) or molecular formula. |
rcsb_search_by_structure | 3D shape-similarity search against a reference PDB assembly or chain. |
rcsb_search_by_seqmotif | Short sequence-motif search (PROSITE pattern, regex, or simple wildcards). |
rcsb_search_strucmotif | 3D structural-motif search: structures sharing a geometric arrangement of specific residues (e.g. a catalytic triad). |
The two text tools (rcsb_search_fulltext, rcsb_search_by_attribute)
also take group_by_identity (100/95/90/70/50/30) to return one representative
per sequence-identity cluster β i.e. non-redundant results. To search
chemical-component attributes, find the path with
rcsb_list_pdb_search_attributes(schema="chemical"), then pass chemical=True to
rcsb_search_by_attribute / rcsb_search_fulltext (usually with return_type="mol_definition").
Both catalogs (structure and chemical) are generated from the live metadata schemas by
scripts/generate_search_attributes.py.
Counting and faceting are output options on every rcsb_search_* tool, not separate
tools: each response includes total_count (the full match count β for "how many ..." run a
search with limit=1 and read it), and passing facets returns a breakdown
(terms/histogram/date_histogram/range/cardinality) instead of hits. The rcsb_search_by_*
service tools (sequence, chemical, structure, seq/struc-motif) also take optional attributes
filters, so e.g. a sequence search can be restricted to an organism in one call.
Sorting is likewise available on every rcsb_search_* tool via sort_by (an
attribute path) + sort_direction (asc/desc), replacing the default score ordering (for
the similarity searches this overrides the similarity-ranked order). Only attributes indexed
for sorting work β those exposing exact_match (strings) or equals (numbers/dates) in
rcsb_list_pdb_search_attributes; sorting is not available for return_type="mol_definition"
(chemical-component results are ranked by score only).
Paging. Every search tool that returns hits accepts limit (1β100, default
10) and offset (default 0). Each response reports total_count, has_more,
and next_offset; to fetch the next page, call the tool again with the same
query and offset set to the returned next_offset.
There is one tool per Data API GraphQL root field. Each takes a list of IDs
(singular lookups = a one-element list) plus an optional fields argument to
override the curated default selection with your own GraphQL sub-selection.
Unknown IDs are reported under not_found. Discover the paths to put in fields
with rcsb_describe_data_object β browse a level, drill into a nested object with
into=, or search the schema by keyword with query= + max_depth=. Every path it
returns is verified against the live schema, so don't guess field names.
| Tool | Object | Example ID |
|---|---|---|
rcsb_get_entries | PDB entries | "4HHB" |
rcsb_get_polymer_entities | Polymer entities (protein/NA) | "4HHB_1" |
rcsb_get_nonpolymer_entities | Ligand/cofactor entities | "4HHB_3" |
rcsb_get_branched_entities | Carbohydrate entities | "5FMB_2" |
rcsb_get_polymer_entity_instances | Polymer chains | "4HHB.A" |
rcsb_get_nonpolymer_entity_instances | Bound-ligand instances | "4HHB.E" |
rcsb_get_branched_entity_instances | Glycan chains | "5FMB.C" |
rcsb_get_assemblies | Biological assemblies | "4HHB-1" |
rcsb_get_interfaces | Assembly interfaces | "1BMV-1.1" |
rcsb_get_chem_comps | Chemical components / ligands | "HEM", "ATP" |
rcsb_get_entry_groups | Entry groups | "G_1002266" |
rcsb_get_polymer_entity_groups | Polymer entity groups (seq. clusters) | "85_70" |
rcsb_get_nonpolymer_entity_groups | Non-polymer entity groups | "ATP" |
rcsb_get_uniprot | UniProt record (single) | "P69905" |
rcsb_get_pubmed | PubMed record (single, integer) | 6726807 |
rcsb_get_group_provenance | Grouping provenance (single) | "provenance_sequence_identity" |
rcsb_describe_data_object | Introspect an object's live GraphQL schema to build a fields= selection: browse a level, drill into a nested object with into=, or search by keyword with query= + max_depth= (flat, incl. nested + cross-object paths). Returns verified dotted paths. The Data API analogue of rcsb_list_pdb_search_attributes. | β |
The Search API only returns identifiers, so a search is the first step: batch the
returned ids into the matching rcsb_get_* tool to fetch titles, organisms, and
other metadata (these tools query the GraphQL endpoint, batching every requested ID
into one request). All 16 typed tools are generated from a single registry in
queries.py (DATA_OBJECTS), so adding a field or
endpoint is a one-line change.
Maps alignments and positional annotations between sequence reference systems
(UNIPROT, NCBI_PROTEIN, NCBI_GENOME, PDB_ENTITY, PDB_INSTANCE). Each
tool takes an optional fields argument to override the default selection; use
rcsb_describe_seqcoord_object to discover what fields are available.
This is the only RCSB API that cross-references NCBI (RefSeq protein /
genome) β the Data API only knows UniProt. So "what NCBI proteins map to a PDB
structure?" is answered by rcsb_seqcoord_alignments, not the Data API. PDB query
ids must be entity-level (4HHB_1), not a bare entry (4HHB); for a whole
entry, query each polymer entity.
| Tool | What it does |
|---|---|
rcsb_seqcoord_alignments | Cross-reference a sequence across PDB / UniProt / NCBI with aligned ranges (e.g. 4HHB_1 β NCBI proteins NP_000508, NP_000549). |
rcsb_seqcoord_annotations | Positional features for one sequence, from one or more annotation sources (UNIPROT, PDB_ENTITY, PDB_INSTANCE, PDB_INTERFACE). |
rcsb_seqcoord_group_alignments | Alignments among members of a sequence group (MATCHING_UNIPROT_ACCESSION / SEQUENCE_IDENTITY). |
rcsb_seqcoord_group_annotations | Annotations across a group; summary=True returns a positional summary. |
rcsb_describe_seqcoord_object | Introspect the live schema to discover fields available on a seqcoord object (for use with fields=). |
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/rcsb-pdb-2)<a href="https://allmcps.com/mcp/rcsb-pdb-2"><img src="https://allmcps.com/api/badge/rcsb-pdb-2?style=directory" alt="RCSB PDB on AllMCPs" /></a>