Get in Touch

Course Outline

Introduction

This section offers a broad overview of when to apply 'machine learning,' outlining key considerations, definitions, advantages, and disadvantages. It covers data types (structured/unstructured/static/streamed), data validity and volume, data-driven versus user-driven analytics, and the distinctions between statistical and machine learning models. It also addresses challenges in unsupervised learning, the bias-variance trade-off, iterative evaluation, cross-validation methods, and supervised/unsupervised/reinforcement learning paradigms.

MAJOR TOPICS

1. Understanding Naive Bayes

  • Core concepts of Bayesian methods
  • Probability theory
  • Joint probability
  • Conditional probability using Bayes' theorem
  • The Naive Bayes algorithm
  • Classification with Naive Bayes
  • The Laplace estimator
  • Handling numeric features in Naive Bayes

2. Understanding Decision Trees

  • The divide and conquer approach
  • The C5.0 decision tree algorithm
  • Selecting the optimal split
  • Pruning decision trees

3. Understanding Neural Networks

  • Transitioning from biological to artificial neurons
  • Activation functions
  • Network architecture
  • Determining the number of layers
  • Direction of information flow
  • Node count per layer
  • Training networks via backpropagation
  • Deep Learning fundamentals

4. Understanding Support Vector Machines

  • Classification via hyperplanes
  • Identifying the maximum margin
  • Linearly separable data scenarios
  • Non-linearly separable data scenarios
  • Applying kernels for non-linear spaces

5. Understanding Clustering

  • Clustering as a machine learning objective
  • The k-means clustering algorithm
  • Utilizing distance for cluster assignment and updates
  • Selecting the optimal number of clusters

6. Measuring Classification Performance

  • Interpreting classification prediction data
  • In-depth analysis of confusion matrices
  • Evaluating performance using confusion matrices
  • Metrics beyond accuracy
  • The kappa statistic
  • Sensitivity and specificity
  • Precision and recall
  • The F-measure
  • Visualizing performance trade-offs
  • ROC curves
  • Predicting future performance
  • The holdout method
  • Cross-validation techniques
  • Bootstrap sampling

7. Optimizing Standard Models for Enhanced Performance

  • Leveraging caret for automated parameter tuning
  • Developing a basic tuned model
  • Customizing the tuning workflow
  • Enhancing model outcomes with meta-learning
  • Concepts of model ensembles
  • Bagging techniques
  • Boosting strategies
  • Random forests
  • Training random forest models
  • Assessing random forest performance

MINOR TOPICS

8. Classification via Nearest Neighbors

  • The kNN algorithm
  • Distance calculation methods
  • Selecting an appropriate k value
  • Data preparation for kNN
  • The lazy nature of the kNN algorithm

9. Classification Rules

  • The separate and conquer strategy
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Understanding Regression

  • Simple linear regression
  • Ordinary least squares estimation
  • Correlation analysis
  • Multiple linear regression

11. Regression and Model Trees

  • Incorporating regression into tree structures

12. Association Rules

  • The Apriori algorithm for rule learning
  • Measuring rule relevance via support and confidence
  • Constructing rule sets using the Apriori principle

Extras

  • Spark/PySpark/MLlib and Multi-armed bandits

Requirements

Python Proficiency

 21 Hours

Number of participants


Price per participant

Testimonials (7)

Upcoming Courses

Related Categories