Sep 13, 2011

New Findings @Boinc!



ENTIRE CATALOG OF FERRET PROTEINS TO DATE



We received a nice comment recently that re-inspired us to get back to work on this site.

This led us, first, straight back to Boinc especially now that we have access to a second, mostly unused computer for processing.
I have to say, after an extended absence, we were very happy to see not one, but TWO new protein related projects @Boinc.

After some research, we decided to add one of the new projects to the existing FerretKnots Boinc team.

Now, in addition to Rosetta@Home we are happy to be participating in

POEM@HOME, sponsored by the Karlsruhe Institute of Technology (KIT) of Germany.

Research Goals of POEM@HOME:
"... a computational approach to
  • predict the biologically active structure of proteins
  • understand the signal-processing mechanisms when the proteins interact with one another
  • understand diseases related to protein malfunction or aggregation
  • develop new drugs on the basis of the three-dimensions structure of biologically important proteins.
The scientific approach behind POEM@HOME is a computational realization of the thermodynamic hypothesis that won C. B. Anfinsen the Nobel Prize in Chemistry in 1972."

Research Goals of Rosetta@Home:
"The goal of our current research is to develop an improved model of intra- and intermolecular interactions and to use this model to predict and design macromolecular structures and interactions.  Prediction and design applications, which can be of great biological interest in their own right, also provide stringent and objective tests that improve the model and increase fundamental understanding.
We use a computer program called Rosetta to carry out protein and design calculations. At the core of Rosetta are potential functions for computing the energies of interactions within and between macromolecules, and methods for finding the lowest energy structure for an amino acid sequence (protein-structure prediction) or a protein-protein complex and for finding the lowest energy amino acid sequence for a protein or protein-protein complex (protein design). ..."

So once again, for the health and welfare of ferrets everywhere, we invite you to participate with us in the FerretKnots Boinc team.

Proteins are Pretty!
(yes, now you can even help by playing a GAME.)


Thanks for stopping in. We hope you'll be back!

All translations copyrighted and owned by myself. All copyrights of their respective owners. No part of this web site may be produced, reproduced, stored in a retrieval system, or transmitted in any form or by any means without the written permission of the copyright owner.

Labels: , , , , , , ,

Jan 27, 2007

2007 week 05: Articles in Proteins

ENTIRE CATALOG OF FERRET PROTEINS TO DATE


Getting one's protein in a bunch -- When quality control fails in cells
Over time, a relatively minor mistake in protein production at the cellular level may lead to serious neurological diseases. But exactly how the cell avoids such mistakes has remained unclear until now. Researchers at Ohio State University found the mechanism that prevents such errors, and explain their findings in the Proceedings of the National Academy of Sciences.

Quantum biology -- Powerful computer models reveal key biological mechanism
Troy, N.Y. -- Using powerful computers to model the intricate dance of atoms and molecules, researchers at Rensselaer Polytechnic Institute have revealed the mechanism behind an important biological reaction. In collaboration with scientists from the Wadsworth Center of the New York State Department of Health, the team is working to harness the reaction to develop a "nanoswitch" for a variety of applications, from targeted drug delivery to genomics and proteomics to sensors.
The research is part of a burgeoning discipline called "quantum biology," which taps the skyrocketing power of today's high-performance computers to precisely model complex biological processes. The secret is quantum mechanics -- the much-touted theory from physics that explains the inherent "weirdness" of the atomic realm.

Microtubule protein interactions visualized en masse
In a new study published online in the open access journal PLoS Biology, Philipp Niethammer, Eric Karsenti, and colleagues investigate the regulation of microtubule dynamics via application of their new method, called visual immunoprecipitation (VIP), which enables simultaneous visualization of multiple protein interactions in cell extracts.

Assignment of polar states for protein amino acid residues using a interaction cluster decomposition algorithm and its application to high resolution protein structure modeling
We have developed a new method (Independent Cluster Decomposition Algorithm, ICDA) for creating all-atom models of proteins given the heavy-atom coordinates, provided by X-ray crystallography, and the pH. In our method the ionization states of titratable residues, the crystallographic mis-assignment of amide orientations in Asn/Gln, and the orientations of OH/SH groups are addressed under the unified framework of polar states assignment. To address the large number of combinatorial possibilities for the polar hydrogen states of the protein, we have devised a novel algorithm to decompose the system into independent interacting clusters, based on the observation of the crucial interdependence between the short range hydrogen bonding network and polar residue states, thus significantly reducing the computational complexity of the problem and making our algorithm tractable using relatively modest computational resources. We utilize an all atom protein force field (OPLS) and a Generalized Born continuum solvation model, in contrast to the various empirical force fields adopted in most previous studies. We have compared our prediction results with a few well-documented methods in the literature (WHATIF, REDUCE). In addition, as a preliminary attempt to couple our polar state assignment method with real structure predictions, we further validate our method using single side chain prediction, which has been demonstrated to be an effective way of validating structure prediction methods without incurring sampling problems. Comparisons of single side chain prediction results after the application of our polar state prediction method with previous results with default polar state assignments indicate a significant improvement in the single side chain predictions for polar residues. Proteins 2007. © 2006 Wiley-Liss, Inc.

Understanding the regulation mechanisms of PAF receptor by agonists and antagonists: Molecular modeling and molecular dynamics simulation studies
Platelet-activating factor receptor (PAFR) is a member of G-protein coupled receptor (GPCR) superfamily. Understanding the regulation mechanisms of PAFR by its agonists and antagonists at the atomic level is essential for designing PAFR antagonists as drug candidates for treating PAF-mediated diseases. In this study, a 3D model of PAFR was constructed by a hierarchical approach integrating homology modeling, molecular docking and molecular dynamics (MD) simulations. Based on the 3D model, regulation mechanisms of PAFR by agonists and antagonists were investigated via three 8-ns MD simulations on the systems of apo-PAFR, PAFR-PAF and PAFR-GB. The simulations revealed that binding of PAF to PAFR triggers the straightening process of the kinked helix VI, leading to its activated state. In contrast, binding of GB to PAFR locks PAFR in its inactive state. Proteins 2007. © 2007 Wiley-Liss, Inc.


Thanks for stopping in! We hope you'll be back!

Ads make the world go around. Help us out!

Labels: , , , , , , , , , , , ,

Jan 2, 2007

2007 week 01: Articles in Folding

ENTIRE CATALOG OF FERRET PROTEINS TO DATE


Steiner minimal trees, twist angles, and the protein folding problem

The Steiner Minimal Tree (SMT) problem determines the minimal length network for connecting a given set of vertices in three-dimensional space. SMTs have been shown to be useful in the geometric modeling and characterization of proteins. Even though the SMT problem is an NP-Hard Optimization problem, one can define planes within the amino acids that have a surprising regularity property for the twist angles of the planes. This angular property is quantified for all amino acids through the Steiner tree topology structure. The twist angle properties and other associated geometric properties unique for the remaining amino acids are documented in this paper. We also examine the relationship between the Steiner ratio [rho] and the torsion energy in amino acids with respect to the side chain torsion angle [chi]1. The [rho] value is shown to be inversely proportional to the torsion energy. Hence, it should be a useful approximation to the potential energy function. Finally, the Steiner ratio is used to evaluate folded and misfolded protein structures. We examine all the native proteins and their decoys at . and compare their Steiner ratio values. Because these decoy structures have been delicately misfolded, they look even more favorable than the native proteins from the potential energy viewpoint. However, the [rho] value of a decoy folded protein is shown to be much closer to the average value of an empirical Steiner ratio for each residue involved than that of the corresponding native one, so that we recognize the native folded structure more easily. The inverse relationship between the Steiner ratio and the energy level in the protein is shown to be a significant measure to distinguish native and decoy structures. These properties should be ultimately useful in the ab initio protein folding prediction. Proteins 2007. © 2006 Wiley-Liss, Inc.

Cooperative folding mechanism of a [beta]-hairpin peptide studied by a multicanonical replica-exchange molecular dynamics simulation

G-peptide is a 16-residue peptide of the C-terminal end of streptococcal protein G B1 domain, which is known to fold into a specific [beta]-hairpin within 6 [mu]s. Here, we study molecular mechanism on the stability and folding of G-peptide by performing a multicanonical replica-exchange (MUCAREM) molecular dynamics simulation with explicit solvent. Unlike the preceding simulations of the same peptide, the simulation was started from an unfolded conformation without any experimental information on the native conformation. In the 278-ns trajectory, we observed three independent folding events. Thus MUCAREM can be estimated to accelerate the folding reaction more than 60 times than the conventional molecular dynamics simulations. The free-energy landscape of the peptide at room temperature shows that there are three essential subevents in the folding pathway to construct the native-like [beta]-hairpin conformation: (i) a hydrophobic collapse of the peptide occurs with the side-chain contacts between Tyr45 and Phe52, (ii) then, the native-like turn is formed accompanying with the hydrogen-bonded network around the turn region, and (iii) finally, the rest of the backbone hydrogen bonds are formed. A number of stable native hydrogen bonds are formed cooperatively during the second stage, suggesting the importance of the formation of the specific turn structure. This is also supported by the accumulation of the nonnative conformations only with the hydrophobic cluster around Tyr45 and Phe52. These simulation results are consistent with high [phis]-values of the turn region observed by experiment. Proteins 2007. © 2006 Wiley-Liss, Inc.

Exploring zipping and assembly as a protein folding principle

It has been proposed that proteins fold by a process called "Zipping and Assembly" (Z&A). Zipping refers to the growth of local substructures within the chain, and assembly refers to the coming together of already-formed pieces. Our interest here is in whether Z&A is a general method that can fold most of sequence space, to global minima, efficiently. Using the HP model, we can address this question by enumerating full conformation and sequence spaces. We find that Z&A reaches the global energy minimum native states, even though it searches only a very small fraction of conformational space, for most sequences in the full sequence space. We find that Z&A, a mechanism-based search, is more efficient in our tests than the replica exchange search method. Folding efficiency is increased for chains having: (a) small loop-closure steps, consistent with observations by Plaxco et al. 1998;277;985-994 that folding rates correlate with contact order, (b) neither too few nor too many nucleation sites per chain, and (c) assembly steps that do not occur too early in the folding process. We find that the efficiency increases with chain length, although our range of chain lengths is limited. We believe these insights may be useful for developing faster protein conformational search algorithms. Proteins 2007. © 2006 Wiley-Liss, Inc.

Strategies for high-throughput comparative modeling: Applications to leverage analysis in structural genomics and protein family organization

The technological breakthroughs in structural genomics were designed to facilitate the solution of a sufficient number of structures, so that as many protein sequences as possible can be structurally characterized with the aid of comparative modeling. The leverage of a solved structure is the number and quality of the models that can be produced using the structure as a template for modeling and may be viewed as the "currency" with which the success of a structural genomics endeavor can be measured. Moreover, the models obtained in this way should be valuable to all biologists. To this end, at the Northeast Structural Genomics Consortium (NESG), a modular computational pipeline for automated high-throughput leverage analysis was devised and used to assess the leverage of the 186 unique NESG structures solved during the first phase of the Protein Structure Initiative (January 2000 to July 2005). Here, the results of this analysis are presented. The number of sequences in the nonredundant protein sequence database covered by quality models produced by the pipeline is [sim]39,000, so that the average leverage is [sim]210 models per structure. Interestingly, only 7900 of these models fulfill the stringent modeling criterion of being at least 30% sequence-identical to the corresponding NESG structures. This study shows how high-throughput modeling increases the efficiency of structure determination efforts by providing enhanced coverage of protein structure space. In addition, the approach is useful in refining the boundaries of structural domains within larger protein sequences, subclassifying sequence diverse protein families, and defining structure-based strategies specific to a particular family. Proteins 2007. © 2006 Wiley-Liss, Inc.

A knowledge-based move set for protein folding

The free energy landscape of protein folding is rugged, occasionally characterized by compact, intermediate states of low free energy. In computational folding, this landscape leads to trapped, compact states with incorrect secondary structure. We devised a residue-specific, protein backbone move set for efficient sampling of protein-like conformations in computational folding simulations. The move set is based on the selection of a small set of backbone dihedral angles, derived from clustering dihedral angles sampled from experimental structures. We show in both simulated annealing and replica exchange Monte Carlo (REMC) simulations that the knowledge-based move set, when compared with a conventional move set, shows statistically significant improved ability at overcoming kinetic barriers, reaching deeper energy minima, and achieving correspondingly lower RMSDs to native structures. The new move set is also more efficient, being able to reach low energy states considerably faster. Use of this move set in determining the energy minimum state and for calculating thermodynamic quantities is discussed. Proteins 2007. © 2006 Wiley-Liss, Inc.


Thanks for stopping in! We hope you'll be back!

Ads make the world go around. Help us out!

Labels: , , , , ,

2007 week 01: Articles in Maths

ENTIRE CATALOG OF FERRET PROTEINS TO DATE


Resources for integrative systems biology: from data through databases to networks and dynamic system models

In systems biology, biologically relevant quantitative modelling of physiological processes requires the integration of experimental data from diverse sources. Recent developments in high-throughput methodologies enable the analysis of the transcriptome, proteome, interactome, metabolome and phenome on a previously unprecedented scale, thus contributing to the deluge of experimental data held in numerous public databases. In this review, we describe some of the databases and simulation tools that are relevant to systems biology and discuss a number of key issues affecting data integration and the challenges these pose to systems-level research.

Strategies for dealing with incomplete information in the modeling of molecular interaction networks

Modelers of molecular interaction networks encounter the paradoxical situation that while large amounts of data are available, these are often insufficient for the formulation and analysis of mathematical models describing the network dynamics. In particular, information on the reaction mechanisms and numerical values of kinetic parameters are usually not available for all but a few well-studied model systems. In this article we review two strategies that have been proposed for dealing with incomplete information in the study of molecular interaction networks: parameter sensitivity analysis and model simplification. These strategies are based on the biologically justified intuition that essential properties of the system dynamics are robust against moderate changes in the value of kinetic parameters or even in the rate laws describing the interactions. Although advanced measurement techniques can be expected to relieve the problem of incomplete information to some extent, the strategies discussed in this article will retain their interest as tools providing an initial characterization of essential properties of the network dynamics.

Dynamic modelling and analysis of biochemical networks: mechanism-based models and model-based experiments

Systems biology applies quantitative, mechanistic modelling to study genetic networks, signal transduction pathways and metabolic networks. Mathematical models of biochemical networks can look very different. An important reason is that the purpose and application of a model are essential for the selection of the best mathematical framework. Fundamental aspects of selecting an appropriate modelling framework and a strategy for model building are discussed.

Concepts and methods from system and control theory provide a sound basis for the further development of improved and dedicated computational tools for systems biology. Identification of the network components and rate constants that are most critical to the output behaviour of the system is one of the major problems raised in systems biology. Current approaches and methods of parameter sensitivity analysis and parameter estimation are reviewed. It is shown how these methods can be applied in the design of model-based experiments which iteratively yield models that are decreasingly wrong and increasingly gain predictive power.

Estimating the parameters of a model for protein-protein interaction graphs

We find accurate approximations for the expected number of three-cycles and unchorded four-cycles under a stochastic distribution for graphs that has been proposed for modelling yeast two-hybrid protein–protein interaction networks. We show that unchorded four-cycles are characteristic motifs under this model and that the count of unchorded four-cycles in the graph is a reliable statistic on which to base parameter estimation. Finally, we test our model against a range of experimental data, obtain parameter estimates from these data and investigate possible improvements in the model. Characterization of this model lays the foundation for its use as a prior distribution in a Bayesian analysis of yeast two-hybrid networks that can potentially aid in identifying false-positive and false-negative results.

A new approach to intensity-dependent normalization of two-channel microarrays

A two-channel microarray measures the relative expression levels of thousands of genes from a pair of biological samples. In order to reliably compare gene expression levels between and within arrays, it is necessary to remove systematic errors that distort the biological signal of interest. The standard for accomplishing this is smoothing "MA-plots" to remove intensity-dependent dye bias and array-specific effects. However, MA methods require strong assumptions, which limit their general applicability. We review these assumptions and derive several practical scenarios in which they fail. The "dye-swap" normalization method has been much less frequently used because it requires two arrays per pair of samples. We show that a dye-swap is accurate under general assumptions, even under intensity-dependent dye bias, and that a dye-swap removes dye bias from a single pair of samples in general. Based on a flexible model of the relationship between mRNA amount and single-channel fluorescence intensity, we demonstrate the general applicability of a dye-swap approach. We then propose a common array dye-swap (CADS) method for the normalization of two-channel microarrays. We show that CADS removes both dye bias and array-specific effects, and preserves the true differential expression signal for every gene under the assumptions of the model.

Regularized linear discriminant analysis and its application in microarrays

In this paper, we introduce a modified version of linear discriminant analysis, called the "shrunken centroids regularized discriminant analysis" (SCRDA). This method generalizes the idea of the "nearest shrunken centroids" (NSC) (Tibshirani and others, 2003) into the classical discriminant analysis. The SCRDA method is specially designed for classification problems in high dimension low sample size situations, for example, microarray data. Through both simulated data and real life data, it is shown that this method performs very well in multivariate classification problems, often outperforms the PAM method (using the NSC algorithm) and can be as competitive as the support vector machines classifiers. It is also suitable for feature elimination purpose and can be used as gene selection method. The open source R package for this method (named "rda") is available on CRAN (http://www.r-project.org) for download and testing.

Are clusters found in one dataset present in another dataset?

In many microarray studies, a cluster defined on one dataset is sought in an independent dataset. If the cluster is found in the new dataset, the cluster is said to be "reproducible" and may be biologically significant. Classifying a new datum to a previously defined cluster can be seen as predicting which of the previously defined clusters is most similar to the new datum. If the new data classified to a cluster are similar, molecularly or clinically, to the data already present in the cluster, then the cluster is reproducible and the corresponding prediction accuracy is high. Here, we take advantage of the connection between reproducibility and prediction accuracy to develop a validation procedure for clusters found in datasets independent of the one in which they were characterized. We define a cluster quality measure called the "in-group proportion" (IGP) and introduce a general procedure for individually validating clusters. Using simulations and real breast cancer datasets, the IGP is compared to four other popular cluster quality measures (homogeneity score, separation score, silhouette width, and weighted average discrepant pairs score). Moreover, simulations and the real breast cancer datasets are used to compare the four versions of the validation procedure which all use the IGP, but differ in the way in which the null distributions are generated. We find that the IGP is the best measure of prediction accuracy, and one version of the validation procedure is the more widely applicable than the other three. An implementation of this algorithm is in a package called "clusterRepro" available through The Comprehensive R Archive Network.

A topologically related singularity suggests a maximum preferred size for protein domains

A variety of protein physicochemical as well as topological properties, demonstrate a scaling behavior relative to chain length. Many of the scalings can be modeled as a power law which is qualitatively similar across the examples. In this article, we suggest a rational explanation to these observations on the basis of both protein connectivity and hydrophobic constraints of residues compactness relative to surface volume. Unexpectedly, in an examination of these relationships, a singularity was shown to exist near 255-270 residues length, and may be associated with an upper limit for domain size. Evaluation of related G-factor data points to a wide range of conformational plasticity near this point. In addition to its theoretical importance, we show by an application of CASP experimental and predicted structures, that the scaling is a practical filter for protein structure prediction. Proteins 2007. © 2006 Wiley-Liss, Inc.

Probabilistic alignment detects remote homologyin a pair of protein sequences without homologous sequence information

Dynamic programming (DP) and its heuristic algorithms are the most fundamental methods for similarity searches of amino acid sequences. Their detection power has been improved by including supplemental information, such as homologous sequences in the profile method. Here, we describe a method, probabilistic alignment (PA), that gives improved detection power, but similarly to the original DP, uses only a pair of amino acid sequences. Receiver operating characteristic (ROC) analysis demonstrated that the PA method is far superior to BLAST, and that its sensitivity and selectivity approach to those of PSI-BLAST. Particularly for orphan proteins having few homologues in the database, PA exhibits much better performance than PSI-BLAST. On the basis of this observation, we applied the PA method to a homology search of two orphan proteins, Latexin and Resuscitation-promoting factor domain. Their molecular functions have been described based on structural similarities, but sequence homologues have not been identified by PSI-BLAST. PA successfully detected sequence homologues for the two proteins and confirmed that the observed structural similarities are the result of an evolutional relationship. Proteins 2007 © 2006 Wiley-Liss, Inc.

iGibbs: Improving Gibbs motif sampler for proteins by sequence clustering and iterative pattern sampling

The motif prediction problem is to predict short, conserved subsequences that are part of a family of sequences, and it is a very important biological problem. Gibbs is one of the first successful motif algorithms and it runs very fast compared with other algorithms, and its search behavior is based on the well-studied Gibbs random sampling. However, motif prediction is a very difficult problem and Gibbs may not predict true motifs in some cases. Thus, the authors explored a possibility of improving the prediction accuracy of Gibbs while retaining its fast runtime performance. In this paper, the authors considered Gibbs only for proteins, not for DNA binding sites. The authors have developed iGibbs, an integrated motif search framework for proteins that employs two previous techniques of their own: one for guiding motif search by clustering sequences and another by pattern refinement. These two techniques are combined to a new double clustering approach to guiding motif search.
...
Tests on the PROSITE database show that their framework improved the prediction accuracy of Gibbs significantly. Compared with more exhaustive search methods like MEME, iGibbs predicted motifs more accurately and runs one order of magnitude faster. Proteins 2007. © 2006 Wiley-Liss, Inc.


Thanks for stopping in! We hope you'll be back!

Ads make the world go around. Help us out!

Labels: , , , , , , , , , ,