The Protein Data Bank (PDB) provides open access to >180,000 biological macromolecules that facilitate understanding of the structure and its function in biology. wwPDB collaboration and expert biocuration is a key to provide the best-curated biological data resource. wwPDB biocurators are at the very front line of structural biology who bridge between the scientists who submit PDB structures and the public who use these data. Our ultimate goal is to support research, training, and education worldwide and facilitate public awareness and understanding of these scientific discoveries. The key to long-term success and sustainability is forming the wwPDB partnership. Currently we have 15 biocurators worldwide working at RCSB PDB, PDBe and PDBj. We maintain a single global archive, no matter which PDB site users download data files from, we ensure identical data files are provided. wwPDB biocurators use a unified OneDep system to curate data. This system provides geographical workload distribution. wwPDB Biocurators setup standard procedures and follow these practices to enforce data standardization among PDB partners. We also review PDB archive and bring the data up to modern standards by conducting remediation effort. Biocuration at wwPDB is important. We manage data deposition, quality, and integrity, and provide integral support to the research community worldwide. To ensure data are Findable, Accessible, Interoperable, and Reusable (FAIR), biocuration at wwPDB focuses on data standardization and quality control. For data standardization, each PDB entry is assigned a unique, persistent identifier, data files are brought to semantic integrity and a common format that can enable large-scale data analysis and distribution. For data quality control, we cross-check with external database references such as UniProt, NCBI, or PubMed. Whatever possible, metadata are conformed to controlled vocabularies and range limits are set to catch data outliers. At the deposition time, biocuration is provided by data depositors on the sequences and taxonomy for polymers and chemical information for ligands if ligands do not match to PDB Chemical Component Dictionary. After data submission, wwPDB biocurators cross-check author’s sequence/taxonomy with UNP/NCBI references, assign existing ligands or add new ligand definitions based on author provided information, provide secondary and quaternary structure as value-added annotation. More importantly, wwPDB biocurators validate structures against experimental data adopting community-recommended validation software and provide depositors an official validation report for their manuscript submission which is now required by many scientific journals. Throughout the COVID-19 pandemic, wwPDB biocuration staff have continued to support research, training and education worldwide. All biocurators have been working remotely since March 2019. The use of unified OneDep system makes this possible with easy transition. We have record high on number of depositions with SARS-CoV-2 structures accounted for 7% of 2020 depositions. For the benefit of global health, as the structural biology community stepped up to provide much-needed insights into SARS-CoV-2, wwPDB biocurators have prioritized SARS-CoV-2 structures and working with depositors closely to release their structures immediately. As of August 1 st 2021, more than 1400 SARS-CoV-2 structures are now freely available from the PDB, reflecting enormous efforts made by the structural biology community in fighting the pandemic. The first structure with PDB ID 6LU7 was deposited on January 26 th 2020 and quickly released in about one week. PDB has recently provided coordinate versioning to allow depositors of the record to update their released structures without change of PDB IDs. Occasionally, rapid PDB data deposition and publication of SARS-CoV-2 structures driven by an understandable sense of urgency has resulted errors in public released PDB structures. The wwPDB coordinate versioning feature has enabled rapid correction of SARS-CoV-2 related structures archived in the PDB. To make the data findable, SARS-CoV-2 structures were updated with standard molecular name, taxonomy, and EC number when UNP references became available.
Read more