Get in Touch
 Duration 21 hours

Course Outline

Comprehensive training curriculum

  1. Introduction to NLP
    • Fundamentals of NLP
    • Popular NLP Frameworks
    • Commercial use cases for NLP
    • Extracting data from the web
    • Retrieving text data via various APIs
    • Managing text corpora, including storage of content and associated metadata
    • Benefits of Python and an NLTK crash course
  2. Practical Approach to Corpora and Datasets
    • The necessity of a corpus
    • Analyzing corpora
    • Categorization of data attributes
    • Various file formats for corpora
    • Preparing datasets for NLP applications
  3. Deconstructing Sentence Structure
    • Core components of NLP
    • Natural language understanding
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Managing ambiguity
  4. Preprocessing Text Data
    • Corpus: Raw Text
      • Sentence tokenization
      • Stemming raw text
      • Lemmatizing raw text
      • Filtering stop words
    • Corpus: Raw Sentences
      • Word tokenization
      • Word lemmatization
    • Utilizing Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Practical and customized preprocessing strategies
  5. Text Data Analysis
    • Essential NLP Features
      • Parsers and parsing techniques
      • POS tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words
    • Statistical NLP Features
      • Linear algebra concepts for NLP
      • Probabilistic theory in NLP
      • TF-IDF
      • Vectorization
      • Encoders and Decoders
      • Normalization
      • Probabilistic Models
    • Advanced Feature Engineering in NLP
      • Word2vec fundamentals
      • Components of the Word2vec model
      • Underlying logic of the Word2vec model
      • Expanding the Word2vec concept
      • Applying the Word2vec model
    • Case Study: Bag of Words application for automatic text summarization using simplified and true Luhn's algorithms
  6. Document Clustering, Classification, and Topic Modeling
    • Document clustering and pattern mining (e.g., hierarchical clustering, k-means)
    • Document comparison and classification using TFIDF, Jaccard, and cosine distance metrics
    • Classifying documents via Naïve Bayes and Maximum Entropy
  7. Identifying Key Text Elements
    • Dimensionality reduction: PCA, SVD, and non-negative matrix factorization
    • Topic modeling and information retrieval via Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis, and Advanced Topic Modeling
    • Measuring sentiment: Positive vs. Negative
    • Item Response Theory
    • Applying Part of Speech tagging to identify people, places, and organizations
    • Advanced topic modeling with Latent Dirichlet Allocation
  9. Case Studies
    • Analyzing unstructured user reviews
    • Sentiment classification and visualization of product review data
    • Extracting usage patterns from search logs
    • Text classification
    • Topic modelling

Requirements

Familiarity with NLP principles and a solid understanding of how AI is applied in business contexts.

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories