Data Science and Its Applications
Book information
Description
The term "data" being mostly used, experimented, analyzed, and researched, "Data Science and its Applications" finds relevance in all domains of research studies including science, engineering, technology, management, mathematics, and many more in wide range of applications such as sentiment analysis, social medial analytics, signal processing, gene analysis, market analysis, healthcare, bioinformatics etc. The book on Data Science and its applications discusses about data science overview, scientific methods, data processing, extraction of meaningful information from data, and insight for developing the concept from different domains, highlighting mathematical and statistical models, operations research, computer programming, machine learning, data visualization, pattern recognition and others. The book also highlights data science implementation and evaluation of performance in several emerging applications such as information retrieval, cognitive science, healthcare, and computer vision. The data analysis covers the role of data science depicting different types of data such as text, image, biomedical signal etc. useful for a wide range of real time applications. The salient features of the book are: Overview, Challenges and Opportunities in Data Science and Real Time Applications Addressing Big Data Issues Useful Machine Learning Methods Disease Detection and Healthcare Applications utilizing Data Science Concepts and Deep Learning Applications in Stock Market, Education, Behavior Analysis, Image Captioning, Gene Analysis and Scene Text Analysis Data Optimization Due to multidisciplinary applications of data science concepts, the book is intended for wide range of readers that include Data Scientists, Big Data Analysists, Research Scholars engaged in Data Science and Machine Learning applications. Cover Half Title Title Page Copyright Page Dedication Table of Contents Preface Acknowledgements Editor Biographies List of Contributors Chapter 1: Introduction to Data Science: Review, Challenges, and Opportunities 1.1 Introduction 1.2 Data Science 1.2.1 Classification 1.2.2 Regression 1.2.3 Deep Learning 1.2.4 Clustering 1.2.5 Association Rules 1.2.6 Times Series Analysis 1.3 Applications of Data Science in Various Domains 1.3.1 Economic Analysis of Electric Consumption 1.3.2 Stock Market Prediction 1.3.3 Bioinformatics 1.3.4 Social Media Analytics 1.3.5 Email Mining 1.3.6 Big Data Analysis Mining Methods 1.4 Challenges and Opportunities 1.4.1 Challenges in Mathematical and Statistical Foundations 1.4.2 Challenges in Social Issues 1.4.3 Data-to-Decision and Actions 1.4.4 Data Storage and Management Systems 1.4.5 Data Quality Enhancement 1.4.6 Deep Analytics and Discovery 1.4.7 High-Performance Processing and Analytics 1.4.8 Networking, Communication, and Interoperation 1.5 Tools for Data Scientists 1.5.1 Cloud Infrastructure 1.5.2 Data/Application Integration 1.5.3 Master Data Management 1.5.4 Data Preparation and Processing 1.5.5 Analytics 1.5.6 Visualization 1.5.7 Programming 1.5.8 High-Performance Processing 1.5.9 Business Intelligence Reporting 1.5.10 Social Network Analysis 1.6 Conclusion References Chapter 2: Recommender Systems: Challenges and Opportunities in the Age of Big Data and Artificial Intelligence 2.1 Introduction 2.2 Methods 2.2.1 Classical 2.2.2 Collaborative Filtering 2.2.3 Content-Based Recommendation 2.2.4 Hybrid FM 2.2.5 Modern Recommender Systems 2.2.6 Data-Driven Recommendations 2.2.7 Knowledge-Driven Recommendations 2.2.8 Cognition-Driven Recommendations 2.3 Application 2.3.1 Classic 2.3.1.1 Multimedia 2.3.1.2 Tourism 2.3.1.3 Food 2.3.1.4 Fashion 2.3.2 Modern 2.3.2.1 Financial Technology (Fintech) 2.3.2.2 Education 2.3.2.3 Recruitment 2.4 Challenges 2.4.1 Cold Start 2.4.2 Context Awareness 2.4.3 Style Awareness 2.5 Advanced Topics 2.5.1 AI-Enabled Recommendations 2.5.2 Cognition Aware 2.5.3 Intelligent Personalization 2.5.4 Intelligent Ranking 2.5.5 Intelligent Customer Engagement 2.6 Conclusion References Chapter 3: Machine Learning for Data Science Applications 3.1 Introduction 3.1.1 Data Science and Machine Learning 3.1.2 Organization of the Chapter 3.2 Linear Regression Models and Issues 3.2.1 Linear Regression 3.2.2 Gradient Descent Method 3.2.3 Regularization 3.2.4 Artificial Neural Networks 3.3 Decision Trees 3.4 Naïve Bayes Model 3.5 Support Vector Machines 3.6 K-Means Clustering Algorithm 3.7 Evaluation of Machine Learning Algorithms 3.7.1 Regression 3.7.1.1 Loss and Error Terms in Linear Regression 3.7.1.2 Correlation Terms in Linear Regression 3.7.2 Classification 3.7.2.1 Binary Classification 3.7.2.2 Multiclass Classification 3.7.3 Clustering 3.8 Some Other Types of Learning Models 3.9 Ensemble Learning with Bagging, Boosting, and Stacking 3.9.1 Bagging 3.9.2 Boosting 3.9.3 Stacking 3.10 Feature Engineering and Dimensionality Reduction 3.11 Applications of Machine Learning in Data Science Bibliography Chapter 4: Classification and Detection of Citrus Diseases Using Deep Learning 4.1 Introduction 4.2 Deep Learning Methods 4.3 Related Work 4.4 Proposed Work 4.4.1 CNN Architecture 4.4.2 CNN Architecture for First Stage Classification 4.4.3 CNN Architecture for Second Stage Classification 4.5 Result and Analysis 4.5.1 Dataset Description 4.5.2 Data Pre- 4.5.3 First- 4.5.4 Second- 4.5.5 Limitations 4.6 Conclusion References Chapter 5: Credibility Assessment of Healthcare Related Social Media Data 5.1 Introduction 5.1.1 User-Generated Content (UGC) 5.1.2 Impact of User-Generated Content 5.1.3 Credibility Assessment of Social Media Data 5.1.4 Credibility Metrics 5.2 Literature Review 5.2.1 Classification of Credibility Assessment of UGC 5.2.2 Credibility Assessment Approaches 5.2.3 Credibility Assessment Platforms 5.3 Credibility Assessment of Healthcare Related Tweets 5.3.1 Data Acquisition and Labeling Using Semi-Supervised Approach (Part I) 5.3.2 Tweet Text Embedding (Part II) 5.3.3 CNN for Text Classification (Part II) 5.4 Experiments and Results 5.5 Conclusion References Chapter 6: Filtering and Spectral Analysis of Time Series Data: A Signal Processing Perspective and Illustrative Application to Stock Market Index Movement Forecasting 6.1 Introduction 6.2 Overview of Digital Filtering Concepts of DSP 6.2.1 Time Domain Representations 6.2.2 Z-Domain Representation 6.2.3 Frequency Domain Representation 6.3 Digital Linear Filtering Analysis with Stochastic Signals 6.4 Time-Series Analysis from Signal Processing Perspective 6.5 Time-Series Models 6.5.1 Partial ACF (PACF) 6.5.2 Akaike Information Criterion (AIC) 6.5.3 Auto-Regressive Integrated Moving Average (ARIMA) Model 6.5.4 Seasonal ARMA (SARMA) and Seasonal ARIMA (SARIMA) Models 6.6 Time-Series Modeling Using Spectal Analysis and Optimum Filtering 6.7 Linear Prediction and Time-Series Modeling 6.8 Forecasting Trends and Seasonality 6.9 Steps for Modeling ARMA ( p, q) Stochastic Process 6.10 Adaptive Filters for Forecasting 6.11 Steps for Forecasting Model Development 6.12 Illustrative Application 6.13 Conclusion and Future Scopes References Chapter 7: Data Science in Education 7.1 Introduction 7.2 Tracing and Identifying Limitations in Current Educational Systems 7.3 Data Science and Learning Analytics: Recent Educational System View 7.4 Learning in the Digital Age: Framework and Features of a SMART Educational System 7.5 Socioeconomic and Technical Challenges in Adopting Learning Analytics in Educational Systems 7.6 Technology Limitations 7.7 Description of a Research-based Pedagogic Innovation for Learning Futures and Underlying Model for IoLT (Internet of Learning Things)–Driven EDM and LA 7.7.1 IoT: What’s in It? 7.7.2 IoT: What It Means to India?: GOI (Government of India) Planning for IoT Expansion in India 7.7.3 GOI Planning for IoT Capacity Building in India 7.7.4 IoT: What It Means to Countries Outside of India, EU as a Representative Group of Countries 7.7.5 Proposed Project Framework: Industry Linked Additive Green Curriculum (ILAGC) 7.7.6 Shift from Internet of Things to Internet of Learning Things 7.8 Emerging Insights Leading to Genesis of ILAGC 7.8.1 Emerging Insight I: Rise of CT, Professional Competitiveness, and Emerging Requirement to Recast Curriculum Implementation to Prepare Students for Learning Skills Futures 7.8.2 Emerging Insight II: Requirement to Recast External Relationships of Curriculum Design and Implementation Necessitates Curriculum Emphasizing Industry Connect and Environment Linkages 7.8.3 Emerging Insight III: With the Rise of CT and Requirement for Industry Connect Emphasizing Curriculum, It Is the Processes of Learning That Are Changing 7.8.4 Emerging Insight IV: Change in Learning Processes Is Leading to the Construct of Additive Curriculum 7.8.5 Summarizing: ILAGC: IoLT Curriculum Design in Brief 7.8.6 How Does Designing Realistic T-L Systems Through Feed-Backward Instruction Design Framework Inform Data Science and Analytics? 7.9 Conclusion References Chapter 8: Spectral Characteristics and Behavioral Analysis of Deep Brain Stimulation by the Nature-Inspired Algorithms 8.1 Introduction 8.2 Related Work 8.3 About Brain Simulation of Alzheimer’s Disease 8.3.1 Perceptual Variations Over Life 8.3.2 Cognitive Domains Differently Decline 8.3.3 Preserved Intelligence Results 8.3.4 Symptoms 8.3.4.1 Conduct 8.3.4.2 Emotional 8.3.4.3 Communal 8.4 Methodology 8.5 Implementation Results 8.6 Conclusion 8.7 Future Enhancement References Chapter 9: Visual Question-Answering System Using Integrated Models of Image Captioning and BERT 9.1 Introduction 9.2 Related Work 9.3 Datasets 9.4 Problem Statement 9.4.1 Proposed Solution 9.5 Methodology 9.5.1 BUTD Captioning (Pythia) 9.5.1.1 The Classifier of Background and Foreground 9.5.1.2 The Regressor of Bounding Box 9.5.1.3 ROI Pooling 9.5.2 Show and Tell: A Neural Image Caption Generator 9.5.3 CaptionBot 9.5.3.1 Computer Vision API 9.5.3.2 Emotion API 9.5.3.3 Bing Image API 9.5.4 Show, Attend, and Tell: Neural Image Caption Generation with Visual Attention 9.5.4.1 Model Details 9.5.4.1.1 Encoder (CNN) 9.5.4.1.2 Decoder (LSTM Network) 9.5.4.2 Learning Stochastic “Hard” versus Deterministic “Soft” Attention 9.5.4.2.1 Stochastic “Hard” Attention 9.5.4.2.2 Deterministic “Soft” Attention 9.5.4.3 Implementation 9.5.5 BERT: Bidirectional Encoder Representations from Transformers 9.6 Experimental Results 9.7 Conclusion Acknowledgments References Chapter 10: Deep Neural Networks for Recommender Systems 10.1 Overview of Recommender Systems 10.2 Jargon Associated with Recommender Systems 10.2.1 Domain 10.2.2 Goal 10.2.3 Context 10.2.4 Personalization 10.2.5 Data 10.2.6 Algorithm 10.3 Building a Deep Neural Network 10.4 Deep Neural Network Architectures 10.4.1 Recurrent Neural Networks (RNN) 10.4.2 Convolutional Neural Network (CNN) 10.5 Recommendation Using Deep Neural Networks 10.6 Tuning Hyperparameters in Deep Neural Networks 10.7 Open Issues in Research 10.8 Conclusion References Chapter 11: Application of Data Science in Supply Chain Management: Real-World Case Study in Logistics 11.1 Introduction 11.2 Related Work 11.3 Smart Supply Chain Management Concept 11.3.1 Demand Forecasting and Stock Optimization 11.3.2 Product Positioning 11.3.3 Picking Zone Prediction 11.3.4 Order Splitting 11.3.5 Order Batching 11.3.6 Order Picking 11.3.7 33D Warehouse Visualization 11.3.8 Smart Data-Driven Transport Management System (TMS) 11.3.8.1 Solving VRP Problems for Instances Divided into Regions (Clusters) 11.3.8.2 Multiphase Adaptive Algorithm for Solving Complex VRP Problems 11.3.8.3 Dynamic Adjustment of the Algorithm Control Parameters 11.3.8.4 Improvements of the Proposed Routes Based on GPS Data 11.3.8.5 Results: Smart Data-Driven Transport Management System 11.3.9 Cluster-Based Improvements to Route-Related Tasks 11.4 Summary Results 11.5 Conclusion Conflict of Interest References Chapter 12: A Case Study on Disease Diagnosis Using Gene Expression Data Classification with Feature Selection: Application of Data Science Techniques in Health Care 12.1 Introduction 12.2 Data Science in Health Care 12.2.1 Drug Discovery 12.2.2 Medical Image Analysis 12.2.3 Predictive Medicine 12.2.4 Medical Record Documentation 12.2.5 Genetics and Genomics 12.2.6 Disease Diagnosis and Prevention 12.2.7 Virtual Assistance to Patients 12.2.8 Research and Clinical Trials 12.3 Issues and Challenges in the Field of Data Science 12.3.1 Inconsistent, Missing, and Inaccurate Data 12.3.2 Imbalanced Data 12.3.3 Cost of Data Collection 12.3.4 Huge Volume of Data 12.3.5 Ethical and Privacy Issues 12.4 Feature Engineering 12.4.1 Feature Extractions 12.4.2 Feature Selection 12.4.2.1 Filter Method 12.4.2.2 Wrapper Method 12.4.2.3 Embedded Method 12.4.3 Feature Weighting 12.4.4 Model Evaluation Metrics 12.5 Case Study of Disease Diagnosis Using Gene Expression Data Classification 12.5.1 Introduction to Gene Expression Dataset 12.5.2 Challenges in Gene Expression Data 12.5.3 Feature Selection and Classification of Gene Expression Data Using Binary Particle Swarm Optimization 12.5.3.1 Binary Particle Swarm Optimization Algorithm 12.5.3.2 Working of Feature Selection Using BPSO 12.5.3.3 Result and Discussion 12.6 Public Dataset and Codes in the Field of Data Science for Health Care 12.7 Conclusion References Chapter 13: Case Studies in Data Optimization Using Python 13.1 Introduction 13.2 Optimization and Data Science 13.3 Literature Review 13.4 Taxonomy of Tools Available for Optimization 13.4.1 Modeling Tools 13.4.2 Solving Tools 13.4.3 Justification for Selecting OR-Tools 13.5 Case Studies Python Prerequisites 13.6 Case Studies: Solving Optimization Problems Through Python 13.6.1 Case Study 1: Product Allocation Problem ( Swarup, Gupta, and Mohan, 2009) 13.6.2 Case Study 2: The Transportation Problem ( Swarup, Gupta, and Mohan, 2009) 13.6.3 Case Study 3: The Assignment Problem ( Swarup, Gupta, and Mohan 2009) 13.7 Conclusions References Chapter 14: Deep Parallel-Embedded BioNER Model for Biomedical Entity Extraction 14.1 Introduction 14.1.1 Rule-Based NER Approach 14.1.2 Directory-Based NER Approach 14.1.3 Machine Learning–Based NER Approach 14.1.4 Deep Learning–Based NER Approach 14.2 Related Work 14.3 Methodology 14.3.1 Character Embedding 14.3.2 Dropout Layer 14.3.3 1D Convolutional Layer 14.3.4 Word Embedding Layer 14.3.5 Casing Embedding Layer 14.3.6 Concatenation Layer 14.3.7 Bidirectional LSTM (BLSTM) 14.3.8 Mathematical Definition 14.3.8.1 Input/Output Layer 14.3.8.2 LSTM Sigmoid and Softmax Layer Softmax Layer 14.4 Evaluation 14.4.1 Corpora Description 14.4.1.1 NCBI 14.4.1.2 JLNPBA 14.4.1.3 BC 4 CHEMD 14.4.2 Competitor System 14.4.3 Evaluation Measure 14.4.4 Hyper-Parameter Setting 14.4.5 Hardware/Software Requirement 14.5 Result and Discussion 14.6 Conclusions and Future Work Conflict of Interest statement Abbreviations Acknowledgment References Chapter 15: Predict the Crime Rate Against Women Using Machine Learning Classification Techniques 15.1 Introduction 15.1.1 Data Analytics 15.1.1.1 Structured Data 15.1.1.2 Unstructured Data 15.1.1.3 Semi-Structured Data 15.2 Machine Learning 15.2.1 Supervised Machine Learning 15.2.2 Unsupervised Machine Learning 15.2.3 Reinforcement Machine Learning 15.3 Data Analytics and Its Classification 15.3.1 Predictive Analytics 15.3.2 Descriptive Analytics 15.3.3 Prescriptive Analytics 15.3.4 Diagnostic Analytics 15.4 Literature Survey 15.5 Methodologies 15.5.1 Data Preprocessing 15.6 Machine Learning Classification Models 15.6.1 Naïve Bayes 15.6.1.1 Drawbacks of Naïve Bayes 15.6.2 Decision Tree Algorithms 15.6.3 K-Nearest Neighbors Model 15.6.4 Support Vector Machine (SVM) 15.6.5 Logistic Regression 15.7 Result and Discussions 15.8 Conclusion References Chapter 16: PageRank–Based Extractive Text Summarization 16.1 Introduction 16.1.1 Types of Text Summarization Process Based on Number of Document 16.1.2 PageRank Algorithm Research Gap 16.2 Related Work 16.3 Proposed Work 16.3.1 Text Summarization Approach 16.3.2 Text Pre-Processing Techniques 16.4 PageRank Methodology Iteration 1: Iteration 2: Iteration 3: Iteration 4: Iteration 5: 16.4.1 PageRank Algorithm: 16.4.2 TextRank Algorithm 16.5 Experimental Setup and Methodology 16.6 Result and Analysis Average Recall Average Precision Average F-Score 16.7 Conclusion References Chapter 17: Scene-Text Analysis 17.1 Introduction 17.2 Literature Survey 17.2.1 Scene-Text Detection 17.2.2 Scene-Text Recognition 17.2.3 Scene-Text Spotting 17.3 Experimental Results 17.3.1 Benchmark Datasets 17.3.2 Performance Metrics 17.3.3 Evaluation for Scene-Text Detection Methods on Benchmark Datasets 17.3.4 Evaluation for Scene-Text Recognition Methods on Benchmark Datasets 17.3.5 Evaluation for Scene-Text Spotting Methods on Benchmark Datasets 17.4 Summary References Index
Similar books
Cognitive Sensors, Volume 2: Applications in Smart Healthcare
2023 · PDF
Cognitive Sensors, Volume 1: Intelligent Sensing, Sensor Data Analysis and Applications
2022 · PDF
Cognitive Sensors, Volume 1: Intelligent sensing, sensor data analysis and applications
2022 · PDF
Statistical Modeling in Machine Learning: Concepts and Applications
2022 · PDF
Advances in Modern Sensors: Physics, design, simulation and applications
2020 · PDF
Modern Optimization Methods for Science, Engineering and Technology
2020 · PDF
Biomedical Signal Processing for Healthcare Applications
2021 · PDF
MySQL® Notes for Professionals book
2018 · PDF