ENGLISH

Processing Metabolomics and Proteomics Data with Open Software: A Practical Guide (ISSN)

Book information

Publisher
Royal Society of Chemistry
Year
2020
ISBN
1788017218, 9781788017213
Language
english
Format
PDF
Filesize
37 MB (38575754 bytes)
Edition
1
Pages
448\460
Time added
2022-02-15 19:19:07

Description

Metabolomics and proteomics allow deep insights into the chemistry and physiological processes of biological systems. These omics methods rely heavily on mass spectrometry, however, building valid models from raw mass spectrometry data is challenging, and the field of data analysis and integration is evolving rapidly. This book will enable researchers, practitioners and students from different backgrounds to analyze metabolomics and proteomics mass spectrometry data. The book contains tutorials, code examples and datasets that facilitate the training and the development of the reader's own programs and workflows. Cover Processing Metabolomics and Proteomics Data with Open Software: A Practical Guide Preface Contents Part A - General Section Chapter 1 - Introduction 1.1 Hypothesis-driven versus Exploratory Research 1.2 Mass Spectrometry Basics 1.2.1 The Sample Introduction Unit 1.2.2 The Separation/Imaging Component 1.2.3 The Ionization Unit 1.2.4 The Mass Analyzer 1.2.5 Fragmentation 1.2.6 Detector 1.2.7 Mass Spectra and Mass Chromatograms 1.2.8 LC-MS Analysis and Data Acquisition Strategies 1.3 Why Open Software for Mass Spectrometry References Chapter 2 - Mass Spectrometry Data Operations and Workflows 2.1 Operations 2.1.1 Formatting 2.1.2 Alignment 2.1.3 Peak Detection 2.1.4 Identification 2.1.5 Calibration 2.1.6 Quantification 2.1.7 Quality Control 2.1.8 Statistical Analysis 2.1.9 Visualization 2.1.10 Deposition 2.2 Workflows References Chapter 3 - Metabolomics 3.1 Introduction to Metabolomics 3.2 Different ‘Flavours’ of Metabolomics 3.3 Technologies for Metabolomics 3.3.1 LC- MS and LC- MS/MS for Metabolomics 3.3.2 GC- MS for Metabolomics 3.3.3 CE- MS for Metabolomics 3.4 LC- MS Processes and Software for Metabolomics 3.4.1 Untargeted LC- MS Metabolomics Tools and Workflows 3.4.1.2 Data Preprocessing and Formula Generation Software 3.4.1.2.1 Construction of EICs.The first step in LC- MS data analysis for metabolomics is the construction of the EICs. This is done to el... 3.4.1.2.2 Peak Picking.The second step in LC- MS analysis involves peak or feature detection from the reconstructed EICs. Peak picking hel... 3.4.1.2.3 Feature Filtering and Grouping.The third and fourth steps in LC- MS data analysis involve feature filtering and feature grouping... 3.4.1.2.4 Feature Alignment.The fifth step involves feature alignment. Metabolomics experiments typically involve the analysis of multiple... 3.4.1.2.5 Statistical Selection.The sixth step in untargeted LC- MS data processing involves the statistical selection of significant feat... 3.4.1.2.6 Mass Matching and Formula Generation.The final step in the untargeted LC- MS metabolomic workflow is feature identification via ... 3.4.1.3 Data Processing for Untargeted LC- MS/MS Data 3.4.2 Targeted LC- MS Metabolomics Tools and Workflows 3.5 GC- MS Metabolomics Tools and Workflows 3.6 CE- MS Metabolomics Workflows and Software 3.6.1 Data Pre- processing Software 3.6.1.1 Peak Picking for CE-MS 3.6.1.2 Feature Filtering and Grouping 3.6.1.3 Feature Alignment 3.6.2 Statistical Analysis 3.6.3 Metabolite Annotation 3.7 Lipidomics Workflows and Software Tools 3.7.1 LC- 3.7.1.1 Lipid Databases and Lipid MS/MS Databases 3.7.2 Shotgun Lipidomics 3.7.3 Imaging Lipidomics of Mass Spectrometry Imaging 3.8 Conclusion References Chapter 4 - Proteomics 4.1 The Proteome: Dimensions, Scales, and Complexity 4.2 Proteomic Experiments and Data Life Cycle 4.3 Signal Processing 4.4 Qualitative Analysis 4.5 Quantitative Analysis 4.6 Getting the Bigger Picture References Chapter 5 - Statistics, Data Mining and Modeling† 5.1 Sample Comparison 5.1.1 Distance Measures 5.1.1.1 Euclidean Distance 5.1.1.2 Manhattan Distance 5.1.1.3 Jaccard Index 5.1.1.4 Pearson Correlation 5.1.1.5 Visualizing Distances in Multiple Samples 5.1.2 Multiple Sample Visualization 5.1.3 Outlier Detection 5.2 Dimensionality Reduction 5.2.1 Principal Component Analysis 5.2.2 Self- organizing Maps 5.3 Cluster Analyses 5.3.1 K- Means 5.3.2 Hierarchical Clustering 5.4 Important Variables 5.4.1 Ranking Peaks 5.4.1.1 Fold- changes 5.4.1.2 t- scores 5.4.2 Biomarker Discovery 5.4.2.1 Working With Peak Intensities 5.4.2.2 Working With Peak Presence/Absence 5.4.2.3 Visualization 5.5 Predictive Models 5.5.1 Machine Learning Introduction 5.5.1.1 Classification Experimental Workflow 5.5.2 Supervised Learning Models 5.5.2.1 Logistic Regression 5.5.2.2 Decision Trees 5.5.2.3 Random Forest 5.5.2.4 Artificial Neural Networks 5.5.2.5 Support Vector Machines 5.5.3 Dataset Partitioning Methods 5.5.3.1 Holdout 5.5.3.2 Cross- validation 5.5.3.3 Bootstrap 5.5.4 Performance Measures 5.5.4.1 ROC Analysis 5.5.5 A Classification Case Study Acknowledgements References Part B - Open MS Programs, Toolkits and Workflow Platforms Chapter 6 - OpenMS and KNIME for Mass Spectrometry Data Processing 6.1 Introduction 6.2 OpenMS for Developers 6.2.1 C++ Library 6.2.2 Data Formats and Raw Data API 6.2.3 Algorithms 6.2.4 TOPP Tools (Developer Perspective) 6.2.5 Visualization 6.2.6 Code Quality and Community 6.2.7 Getting Started with the OpenMS Library for Developers 6.3 OpenMS for Users 6.3.1 TOPP Tools (User Perspective) 6.3.2 Getting Started with OpenMS for Users 6.3.3 Workflows in MS 6.3.3.1 OpenMS in KNIME 6.3.3.2 Getting Started with OpenMS in KNIME 6.3.3.3 Other Workflow Systems with OpenMS Integrations 6.3.4 Peptide Identification and Protein Inference 6.3.4.1 Search Engine Choice 6.3.4.2 Sequence Database 6.3.4.3 Identification Post- processing 6.3.4.4 Protein Inference 6.3.5 Further Peptide Identification Methods 6.3.5.1 De Novo Peptide Search 6.3.5.2 Spectral Library Search 6.3.6 Additional Supported Methods 6.3.6.1 Phosphosite Localization 6.3.6.2 Spectral Clustering 6.3.7 Peptide and Protein Quantification 6.3.7.1 Feature Finding 6.3.7.2 Combining Post- processed Identification with Quantification Data (Peptide Level) 6.3.7.3 Retention Time Alignment, Feature Linking and Generation of Peptide Level Results 6.3.7.4 Generate Protein Level Results 6.3.8 Additional Supported Quantification Methods 6.3.8.1 Label Free Quantification Based on Identification Data 6.3.8.2 Quantification Using Chemical or Isotopic Labeling 6.3.8.3 Quantification Using Isobaric Labeling 6.3.9 Targeted Analysis 6.3.10 Metabolomics 6.3.10.1 Metabolite Quantification 6.3.10.2 Metabolite Identification 6.3.10.2.1 Compound Databases.The tool AccurateMassSearch is the first step towards compound identification and can be used to annotate det... 6.3.10.2.2 Spectral Library Search.Searching the accurate mass of unidentified metabolites against a compound database will provide insight... 6.3.10.2.3 De Novo.De novo identification has the advantage that it does not rely on a spectral library of previously measured compounds. S... 6.3.11 Metaproteomics 6.3.12 Cross- linking MS 6.3.13 RNA (Modification) Analysis 6.3.14 Visualization Capabilities (User Perspective) 6.3.15 Containerization and Reproducibility Acknowledgements References Chapter 7 - Metabolomics Data Analysis Using MZmine 7.1 Introduction 7.2 Feature Detection 7.2.1 ADAP Feature Detection Methods 7.2.2 GridMass – 2D Feature Detection 7.2.3 Evaluation of Feature Detection Methods 7.3 Spectral Deconvolution 7.3.1 Hierarchical Clustering Method 7.3.2 MCR Method 7.4 Compound Identification 7.4.1 Chemical Formula Prediction 7.4.2 Compound Database Search (MS1 Level Identification) 7.4.3 Machine- learning- based Structure Prediction (MS/MS Level Identification) 7.4.4 Spectral Similarity 7.4.5 Lipid Identification 7.5 Batch Mode 7.6 Conclusions Acknowledgements References Chapter 8 - Pre-processing and Analysis of Metabolomics Data with XCMS/R and XCMS Online† 8.1 Introduction 8.2 Example Project: A Biological Background and Analytical Question 8.3 Raw Data Preparation and Preview 8.3.1 Data and Code Availability 8.3.2 Converting Raw Files in Vendor Format 8.3.3 mzML File Preview 8.4 XCMS/R 8.4.1 Directory Structure of Data 8.4.2 RStudio Editor for R 8.4.3 Installing the R Packages 8.4.4 Loading and Running the R Script 8.4.5 Loading Required R Libraries 8.4.6 Reading and Annotating Raw Data 8.4.7 Defining Colours 8.4.8 Reading the Raw Data 8.4.9 Plotting Base Peak Chromatograms 8.4.10 Total Ion Current Box Plot 8.4.11 Test a Single Feature 8.4.12 Feature Detection 8.4.13 Retention Time Correction 8.4.14 Grouping/Binning Features 8.4.15 Filling Data for Missing Peaks 8.4.16 Creating a Feature Summary 8.4.17 Histogram of Features 8.4.18 Sub- setting and Exporting Features 8.4.19 Principal Component Analysis (PCA) 8.4.20 Hierarchical Cluster Analysis (HCA) 8.4.21 PAM Clustering 8.4.22 Additional Data Mining with Rattle 8.4.23 Revision the Data Processing History 8.4.24 Further Statistical Evaluation 8.4.25 Running the Complete Workflow 8.4.26 Metabolite Identification 8.5 XCMS Online 8.5.1 Registration as a User 8.5.2 Dataset Specification 8.5.3 Preparing Data for Uploading 8.5.4 Uploading Datasets 8.5.5 Available Analysis Types 8.5.6 Creating and Submitting a Job 8.5.7 Visualizing the Results 8.5.8 Non- multi- dimensional Scaling (NMDS) Analysis 8.5.9 Metabolic Cloud Plot 8.5.10 Pathway Cloud Plot and Systems Biology Results 8.5.11 Interpreting Results 8.5.12 Data Sharing 8.6 Comparing XCMS/R and XCMS Online References Chapter 9 - Statistical Evaluation and Integration of Multi- omics Data with MetaboAnalyst 9.1 Introduction to MetaboAnalyst 9.2 MetaboAnalyst Overview 9.3 Data Formats and Data Requirements 9.4 General Statistical Analysis with MetaboAnalyst 9.5 Enrichment Analysis and Pathway Analysis 9.6 MS Peaks- to- Pathways and Mummichog 9.7 Summary References Chapter 10 - Modular metaX Pipeline for Processing Untargeted Metabolomics Data† 10.1 Introduction 10.2 Data Preparation 10.3 Data Pre- processing 10.4 Data Quality Assessment 10.5 Normalization Evaluation 10.6 Other Functions in metaX 10.7 Integrated Function metaXpipe 10.8 Applications Acknowledgements References Chapter 11 - Metabolite Annotation With CEU Mass Mediator 11.1 Introduction 11.2 Non- maintained Resources 11.3 CEU Mass Mediator Acknowledgements References Chapter 12 - Metabolite Annotation Using In Silico Generated Compounds: MINE and BioTransformer 12.1 Introduction 12.2 Non- maintained Resources 12.3 Metabolic In Silico Network Expansion Databases 12.3.1 Structure Search 12.3.2 MS Adduct Search 12.3.3 MS/MS Search 12.3.4 Compound Page 12.4 BioTransformer 12.4.1 The BioTransformer Metabolism Prediction Tool (BMPT) 12.4.2 The BioTransformer Metabolism Identification Tool (BMIT) Acknowledgements References Chapter 13 - Trans-Proteomic Pipeline for the Identification, Validation, and Quantification of Proteins† 13.1 Introduction 13.2 Using the Tools 13.3 Conclusion Acknowledgements References Chapter 14 - Quantitative Proteomics Data Analysis with PANDA, LFAQ and PANDA-view 14.1 Introduction 14.2 Materials 14.2.1 Hardware Requirements 14.2.2 Software Requirements 14.2.3 Software Installation 14.3 Procedures 14.3.1 Relative Protein Quantification Using PANDA 14.3.1.1 Set Parameters in the Data Interface 14.3.1.2 Set Parameters in the Parameter Interface 14.3.1.3 Set Parameters in the Progress Interface 14.3.1.4 Anticipated Results 14.3.2 Absolute Protein Quantification Using LFAQ 14.3.2.1 Set Parameters and Generate a New Parameter File in GUI Set Parameters About the Input Data Set Parameters for Quantification 14.3.2.2 Load an Existing Parameter File 14.3.2.3 Run LFAQ 14.3.2.4 Anticipated Results 14.3.2.4.1 Annotation of ProteinResultsExperimentName.txt.This file contains the quantification results of proteins from the experiment Exp... 14.3.2.4.2 Annotation of ProteinMergedResults.txt.This file contains the merged protein quantification results of LFAQ from different exper... 14.3.2.4.3 Annotation of LFAQResultsForStandardProteinsExperimentName.txt.This file contains the actual abundances and the predicted abunda... 14.3.3 Post- processes of the Quantification Results Using PANDA- view 14.3.3.1 Missing Value Imputation and Normalization 14.3.3.2 Statistical Analysis of Quantification Result 14.3.3.3 Visualization of Quantification Results 14.4 Discussions and Conclusions Acknowledgements References Chapter 15 - Proteomic Workflows with R/ R Markdown 15.1 Introduction 15.1.1 R 15.1.2 Markdown and R Markdown 15.2 Materials, Results, and Discussion 15.2.1 Example of a Simple Workflow 15.2.1.1 Load R Packages and Specify the Search Parameters and Sequence Database 15.2.1.2 Convert Raw Data to mzML 15.2.1.3 Perform Database Search 15.2.1.4 Validate Search Results With PeptideProphet11 15.2.1.5 Protein Inference using ProteinProphet12 15.2.1.6 Collect iTRAQ- quantitation from the Quantitation.tsv File and Make Violin Plot 15.2.1.7 Summary of the Simple Workflow 15.2.2 Example of an Advanced Workflow 15.2.2.1 Methods 15.2.2.2 Results 15.2.2.3 Summary of the Advanced Workflow 15.3 Conclusion Acknowledgements References Chapter 16 - Python in Proteomics 16.1 Installation 16.2 Getting Started 16.3 Plotting 16.4 Chemistry 16.4.1 Elements 16.4.2 Molecular Formula 16.4.3 Isotopic Distributions 16.4.4 Amino Acids 16.5 Peptides and Proteins 16.5.1 Amino Acid Sequences 16.5.2 Molecular Formula 16.5.3 Modified Sequences 16.5.4 Proteins 16.5.5 TheoreticalSpectrumGenerator 16.6 Digestion 16.6.1 Proteolytic Digestion with Trypsin 16.6.2 Proteolytic Digestion with Lys- C 16.7 Simple Data Manipulation 16.7.1 Filtering Spectra 16.7.2 Filtering by MS Level 16.7.3 Filtering by Scan Number 16.7.4 Filtering Spectra and Peaks 16.7.5 Memory Management 16.8 Example: Peptide Search 16.9 pyOpenMS in R 16.9.1 Install the “reticulate” R Package 16.9.2 Import Pyopenms in R References Chapter 17 - Mass Spectrometry Development Kit (MSDK): a Java Library for Mass Spectrometry Data Processing 17.1 Introduction 17.2 Architecture 17.3 MSDK Modules 17.3.1 Raw Data File Format Support 17.3.2 Basic Spectra Processing Algorithms 17.3.3 Feature Detection 17.3.4 Compound Identification 17.4 Future Plans 17.5 Conclusions Acknowledgements References Chapter 18 - MASSyPup64: Linux Live System for Mass Spectrometry Data Processing 18.1 Introduction 18.2 ‘Installation- free’ MS Data Processing 18.3 Programs Installed on MASSyPup64 18.4 Preparing and Starting MASSyPup64 18.5 Example: Determination of Protein Weight with ESIprot 18.6 Remastering the Live USB 18.7 Future of MASSyPup64 References Chapter 19 - Cross- platform Software Development and Distribution with Bioconda and BioContainers 19.1 Introduction 19.2 From Tools to Bioconda Packages 19.3 Bioinformatics Containers 19.4 Containers Deployment and Workflows 19.5 Conclusions References Part C - Conclusion Chapter 20 - Concluding Remarks and Perspectives References Subject Index

Similar books

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

2010 · PDF

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

1858 · PDF

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

2022 · PDF

The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians

The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians

1809 · PDF