Crystal structure of a novel non‐Pfam protein AF1514 from <i>Archeoglobus fulgidus</i> DSM 4304 solved by S‐SAD using a Cr X‐ray source
Yang Li, Pazilat Bahti, Neil Shaw, Gaojie Song, Shunmei Chen, Xuejun Zhang, Min Zhang, Chongyun Cheng, Jie Yin, J. Zhu, Hua Zhang, Dongsheng Che, Hao Xu, Abdulla Abbas, Bi‐Cheng Wang, Zhi‐Jie Liu
- Year
- 2008
- Citations
- 8
- Access
- Open access
Abstract
Many computational tools have been developed recently to accurately predict the structure of a protein from its amino acid sequence.1-3 In general, if a query protein sequence shares at least 30% sequence identity with a protein sequence whose 3D structure has been determined, the structure of this query sequence can be modeled based on the template structure, using MODELLER software for example.3, 4 Computational software, however, cannot guarantee accurate prediction for those new proteins that share low sequence similarity in PDB. Therefore, experimental methods such as X-ray crystallography and nuclear magnetic resonance (NMR) are still the main approaches for a protein structural study.5, 6 The current target selection strategy of most structural genomics centers7-9 mainly focuses on the representatives of manually curated protein families (Pfam),10-12 that is, the selected protein sequence shares at least one conserved domain with other members within a family. In this way, the solved representative structures can be used as structural templates to predict structures of the remaining protein sequences in the same family using computational tools. It has been shown that this “Pfam” target selection strategy increases not only the number of novel structures, but also the number of new folds.13 However, over-emphasis of Pfam and ignoring non-Pfam sequences (i.e., not sharing any conserved domain in Pfam) in target selection might lead to biased distribution of Pfam and non-Pfam structures in PDB, and possibly slow the growth rate of new structures and folds. Our analysis on 150 microbial genomes showed that non-Pfam sequences account for 25–30% of all Open Reading Frames (ORFs) for most genomes, and some could reach to 60% (unpublished data). The high percentage of non-Pfams over all ORFs reminds us non-Pfams should not be neglected while devising a target selection strategy. On the other hand, these non-Pfam sequences for each genome are either paralogous non-Pfam (in which sequences have homologous partners within the same organism), or orthologous non-Pfam (in which sequences have orthologous partners in the closely related organisms), or singleton non-Pfam (in which sequences have only one copy in the organism). These three non-Pfam kinds are either organism-specific or genus-specific, implying the possible existence of undiscovered unique features of non-Pfams, such as unique SCOP fold or CATH topology. Many Pfam sequences have significant biological meaning, while the functions of most non-Pfam sequences are unknown to date. However, this does not mean that non-Pfam sequences are biologically less meaningful. Some non-Pfam sequences and structures are predicted to play important roles for the uniqueness of these organisms. Therefore, a non-Pfam selection strategy will not only accelerate the expansion of SCOP fold space, but also help biologists understand the functions of these proteins based on structures. The present work describes the structure of AF1514, a non-Pfam protein from Archeoglobus fulgidus with unknown function, solved at 1.8 Å resolution by using anomalous signal of sulfur generated by chromium X rays (wavelength = 2.29 Å). E coli BL21 was freshly transformed with plasmid containing AF1514 gene. Cells were grown at 37°C until culture density reached OD600 nm − 0.8. The culture was cooled down to 12°C and induced with 0.2 mM IPTG for 40 h. Cells were harvested by centrifugation and lysed by sonication. Cell debris was removed by centrifugation and the clarified supernatant subjected to Ni-affinity chromatography. The protein was further purified using size exclusion chromatography. The purified AF1514 protein was divided into two equal aliquots. One aliquot was methylated as described previously,14-16 while the other aliquot was directly concentrated without any chemical modification. Both, methylated and non-methylated protein samples were concentrated to ∼18 mg mL−1 in 20 mM Tris-HCl, pH 8.0, 200 mM N
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
Genetic Programming: On the Programming of Computers by Means of Natural Selection
John R. Koza
1992