2026 Batch Open
Course 2 — MLOps & Production AI Engineering (60-Class Career Track)
MLOPS-60 • 7 Months Career Track
Cloud, DevOps & MLOps MLOPS-60 Lab & Live Online 60 Classes (120 Hours)

Course 2 — MLOps & Production AI Engineering (60-Class Career Track)

Goal: Take engineers from machine learning models to production scale. Master model serving APIs, experiment tracking, data versioning, feature stores, automated continuous training (CT), model monitoring, and scalable AI infrastructure on Kubernetes.

Core Engineering Competencies

Module 1: High-performance ML serving with FastAPI, Pydantic validation, async endpoints, and Docker containerization
Module 2: Experiment tracking, parameter logging, and enterprise model registry with MLflow on AWS
Module 3: Data version control (DVC), Great Expectations validation, and Feast feature store architecture
Module 4: Automated CI/CD/CT pipelines, Continuous Machine Learning (CML), and automated retraining triggers
Module 5: Scalable ML deployment on Kubernetes, TorchServe/Triton, KServe serverless inference, and Kubeflow pipelines
Module 6: Production model monitoring, data drift / concept drift detection with Evidently AI, and final MLOps capstone
Lab: ₹2,000 ₹1,500
Live Online: ₹2,000 ₹1,500
₹1,500/mo or 10% upfront discount
WEEKEND BATCH
Schedule: Saturday & Sunday
Class Duration: 2 Hours / Class
Lab Ratio: 40% Concept / 60% Lab
Location: Bagnan Main Campus
Curriculum Breakdown

Detailed Class-by-Class Syllabus

Total: 60 Classes (120 Hours)

What is MLOps, ML technical debt, Python for MLOps, FastAPI model serving, Pydantic request validation, model serialization (ONNX/Joblib), and testing.

Class 1: What is MLOps? 120m

DevOps vs MLOps, why 85% of ML models fail to reach production, MLOps pillars (Code, Data, Model).

Class 2: ML Technical Debt & Production Gaps 120m

Hidden technical debt in ML systems, glue code, pipeline jungles, dead experimental code.

Class 3: Modern MLOps Stack Overview 120m

Architecture blueprint: Data Ingestion → Feature Store → Training → Registry → Serving → Monitoring.

Class 4: Python for MLOps & Type Hinting 120m

Clean code principles, type hints (typing module), dataclasses, Pydantic validation schemas.

Class 5: FastAPI Model Serving 120m

Building REST APIs with FastAPI, routing, path/query parameters, dependency injection for model loading.

Class 6: Request Validation & Error Handling 120m

Pydantic BaseModel for feature inputs, input boundary validation, HTTP error status handling.

Class 7: Model Serialization & Formats 120m

Pickle vs Joblib vs ONNX (Open Neural Network Exchange) vs TorchScript for portable inference.

Class 8: High-Performance & Batch Inference 120m

Async endpoints, batch prediction vectors, minimizing latency and CPU overhead.

Class 9: Testing ML Applications 120m

Unit testing with Pytest, testing API endpoints with TestClient, data schema validation tests.

Class 10: FastAPI ML Microservice Project 120m

Containerizing and deploying a production-ready ML inference microservice with Docker and Swagger documentation.

Level Capstone Project 100 Marks

Project: Production-Ready FastAPI Machine Learning Microservice

Trained ML Model → ONNX Export → FastAPI Asynchronous Endpoint → Pydantic Schema Validation → Multi-Stage Docker Container.

Python FastAPI Pydantic ONNX Docker Pytest

Reproducibility, MLflow Tracking, autologging, artifacts, metric logging, remote MLflow server on AWS, and Model Registry workflows.

Class 11: The Reproducibility Problem in ML 120m

Why model training must be version-controlled, tracking hyperparameters, metrics, and seeds.

Class 12: MLflow Architecture 120m

MLflow Tracking, MLflow Projects, MLflow Models, MLflow Model Registry core components.

Class 13: Logging Parameters & Metrics 120m

mlflow.log_param(), mlflow.log_metric(), mlflow.log_artifact(), tracking loss curves over training epochs.

Class 14: MLflow Autologging 120m

Autologging with Scikit-learn, PyTorch, TensorFlow, and XGBoost with zero boilerplate code.

Class 15: Remote MLflow Tracking Server 120m

Setting up remote MLflow backend on AWS EC2, PostgreSQL backend store, and S3 artifact store.

Class 16: MLflow Model Registry 120m

Registering models, semantic versioning (v1, v2, v3), model stages: Staging, Production, Archived.

Class 17: Automated Model Promotion 120m

Automating model stage transitions based on validation metric thresholds (e.g. F1-score > 0.90).

Class 18: Weights & Biases (W&B) 120m

Overview of W&B dashboards, hyperparameter sweeps, and collaborative model evaluation.

Class 19: MLflow Models Packaging 120m

Packaging models with MLmodel flavor definitions for universal deployment targets.

Class 20: Enterprise Model Registry Project 120m

Setting up production MLflow server on AWS, running hyperparameter experiments, and promoting best model to Production.

Level Capstone Project 100 Marks

Project: Centralized Enterprise Experiment Tracking & Model Registry

Remote MLflow on AWS EC2 + S3 Artifacts + RDS PostgreSQL → Hyperparameter Experimentation → Automated Validation Gate → Production Promotion.

MLflow AWS S3 AWS RDS PostgreSQL Scikit-Learn PyTorch

Data Version Control (DVC), S3 remote storage, reproducible pipelines, Great Expectations data validation, and Feast feature store.

Class 21: Data Version Control (DVC) Fundamentals 120m

Why Git cannot handle large datasets, pointer files (.dvc), content-addressable storage.

Class 22: DVC with Cloud Remotes (AWS S3) 120m

Configuring S3 remote storage in DVC, dvc add, dvc push, dvc pull across developer workstations.

Class 23: Building Reproducible DVC Pipelines 120m

Writing dvc.yaml pipelines, stage dependencies, output caching, executing end-to-end with dvc repro.

Class 24: Data Validation at Scale 120m

Data testing principles, detecting silent data corruption and unexpected schema shifts.

Class 25: Great Expectations Framework 120m

Defining Expectation Suites, automated data profiling, generating HTML Data Docs.

Class 26: Introduction to Feature Stores 120m

What is a feature store? Eliminating training-serving skew, feature reusability across models.

Class 27: Feast Feature Store Architecture 120m

Feast definitions, Entities, FeatureViews, source datasets, feature repository structure.

Class 28: Offline vs Online Feature Stores 120m

Offline store (Parquet/Snowflake for batch training) vs Online store (Redis for low-latency inference).

Class 29: Serving Real-time Features 120m

Materializing features to Redis, fetching real-time feature vectors during live model inference.

Class 30: Data Versioning & Feature Store Project 120m

Building an end-to-end data pipeline with DVC, Great Expectations validation, and Feast feature serving.

Level Capstone Project 100 Marks

Project: Versioned Data Pipeline & Feast Feature Store

Raw Data → DVC + S3 Versioning → Great Expectations Data Validation → Feast Feature Store → Real-time Redis Serving.

DVC AWS S3 Great Expectations Feast Redis Python

MLOps maturity levels, Docker for ML, GitHub Actions CI for ML, Continuous Machine Learning (CML), automated model training, and evaluation gates.

Class 31: MLOps Maturity Levels 120m

Level 0: Manual process, Level 1: ML Pipeline automation, Level 2: CI/CD/CT automated pipeline.

Class 32: Docker for ML Workloads 120m

Optimizing Docker images for ML, CUDA GPU support, lightweight slim base images.

Class 33: Multi-Stage Dockerfiles for PyTorch 120m

Building production container images with separate build and runtime environments.

Class 34: GitHub Actions CI for ML 120m

Automated linting (Flake8, Black), Pytest test suites, and data validation in GitHub Actions.

Class 35: Continuous Machine Learning (CML) 120m

Automating pull-request model reports, markdown metrics tables, ROC curve image generation in PR comments.

Class 36: Automated Model Training (Continuous Training) 120m

Triggering automated training runs on new data ingestion or schedule via GitHub Actions.

Class 37: Model Evaluation Quality Gates 120m

Automated metric comparison: comparing newly trained candidate model against active production champion model.

Class 38: Automated Container Build & ECR Push 120m

Building Dockerized inference service upon successful model training and pushing to AWS ECR.

Class 39: Zero-Downtime Model Deployment 120m

Automating model artifact sync to production servers without service interruption.

Class 40: End-to-End CI/CD/CT Pipeline Project 120m

Code/Data Commit → Automated Training → Quality Gate Check → Docker Build → Automated Staging Deployment.

Level Capstone Project 100 Marks

Project: Automated CI/CD/CT Pipeline for Machine Learning

Git Commit → GitHub Actions CI → DVC Data Pull → Automated Training → Evaluation Quality Gate → Container Registry → Deployment.

GitHub Actions CML DVC Docker AWS ECR Python

Containerized ML serving on Kubernetes, TorchServe/Triton, KServe serverless ML, Kubeflow Pipelines (KFP), Canary rollouts, and Autoscaling.

Class 41: Containerized ML Serving on Kubernetes 120m

Deploying FastAPI ML inference Pods, Services, and Nginx Ingress on Kubernetes.

Class 42: Dedicated Model Servers (TorchServe / Triton) 120m

NVIDIA Triton Inference Server architecture, dynamic batching, multi-model execution.

Class 43: KServe Serverless Inference on K8s 120m

KServe (KFServing) CRDs, scale-to-zero serverless inference, inference graphs.

Class 44: Kubeflow Fundamentals 120m

Kubeflow architecture on Kubernetes, Kubeflow dashboard, Central Dashboard components.

Class 45: Kubeflow Pipelines (KFP) 120m

Writing Kubeflow pipeline components in Python with @component decorator, passing data artifacts.

Class 46: Orchestrating End-to-End Pipelines 120m

Compiling and submitting Kubeflow training workflows with data preprocessing, training, and evaluation.

Class 47: Deployment Strategies: Canary & A/B Testing 120m

Routing 10% traffic to candidate model and 90% to champion model on Kubernetes.

Class 48: Shadow Deployments & Blue-Green Rollouts 120m

Mirroring live production inference requests to test new models without impacting users.

Class 49: Autoscaling ML Workloads (HPA) 120m

Horizontal Pod Autoscaler configuration based on GPU utilization and inference request metrics.

Class 50: Kubernetes ML Serving Fleet Project 120m

Deploying a high-throughput multi-model inference fleet on AWS EKS with Canary rollout and Autoscaling.

Level Capstone Project 100 Marks

Project: Scalable Multi-Model Serving Fleet on AWS EKS

AWS EKS Kubernetes Cluster → KServe / FastAPI Serving → Ingress Canary Traffic Splitting (90/10) → Horizontal Pod Autoscaling.

Kubernetes AWS EKS KServe FastAPI Helm Docker

Production ML monitoring, data drift, concept drift, statistical drift tests (KS, PSI), Evidently AI, Prometheus/Grafana, LLMOps, and Final Capstone.

Class 51: Production ML Monitoring 120m

Operational metrics (latency, error rate) vs Data Science metrics (drift, feature stability, accuracy).

Class 52: Data Drift vs Concept Drift 120m

Why production models degrade: distribution changes, seasonal trends, covariate shift.

Class 53: Statistical Drift Detection Methods 120m

Kolmogorov-Smirnov (KS) test, Population Stability Index (PSI), Wasserstein Distance.

Class 54: Evidently AI Framework 120m

Automated data drift dashboards, target drift reports, data quality test suites.

Class 55: Prometheus & Grafana for ML Metrics 120m

Exporting drift scores and prediction metrics to Prometheus, real-time Grafana dashboards.

Class 56: Automated Retraining on Drift Alert 120m

Triggering automated retraining pipelines in GitHub Actions/Kubeflow upon drift threshold breaches.

Class 57: LLMOps Fundamentals 120m

Monitoring Large Language Models, prompt drift, token usage, latency, hallucination detection.

Class 58: Production AI Security & Governance 120m

Model explainability (SHAP), fairness, audit logging, model lineage tracking.

Class 59: Final Capstone Development 120m

Building the end-to-end production MLOps system: Data Ingestion → DVC → MLflow → CI/CD/CT → EKS Serving → Evidently AI Drift Monitoring.

Class 60: Final Presentation + Assessment 120m

1-on-1 viva defense, live production architecture demonstration, and Course 2 MLOps Certification.

Level Capstone Project 100 Marks

Course 2 Capstone: Production End-to-End MLOps & Continuous Training Platform on AWS & Kubernetes

Data Versioning (DVC + S3) → Experiment Tracking (MLflow) → Automated CI/CD/CT (GitHub Actions) → Kubernetes Deployment (AWS EKS) → Real-Time Monitoring & Drift Detection (Evidently AI + Prometheus).

FastAPI MLflow DVC GitHub Actions Kubernetes AWS EKS Evidently AI Prometheus Grafana Feast