Project Overview
This project involved setting up a cloud monitoring solution for AWS and Azure instances using Prometheus, Telegraf, and Grafana. The Prometheus servers were deployed automatically via Terraform on AWS, while Telegraf agents were dynamically deployed on both AWS EC2 instances and Azure VMs using Ansible roles. Telegraf was automatically installed during the creation of new instances via Terraform, ensuring all instances were continuously monitored. Alerts were configured with Alertmanager, triggering PagerDuty notifications integrated with Microsoft Teams for real-time incident responses. The entire setup was secured with a site-to-site VPN between AWS and Azure, ensuring seamless monitoring of both environments.
Technologies and Services Used
Terraform: Automated the deployment of Prometheus servers and Telegraf on AWS and Azure.
Ansible: Managed Telegraf deployment automatically on AWS EC2 instances and Azure VMs during provisioning via Terraform.
Prometheus: Deployed on AWS for storing and querying system metrics, with automatic deployment using Terraform.
Telegraf: Automatically installed and configured using Ansible, collecting metrics from instances and VMs.
Grafana: Used for real-time visualization of metrics, with customizable dashboards for monitoring both cloud environments.
Alertmanager: Managed alerts based on Prometheus thresholds, routing them to PagerDuty for incident management.
PagerDuty: Integrated for incident escalation, connected to Microsoft Teams for notifications.
Microsoft Teams: Used for real-time notifications through integration with PagerDuty.
AWS and Azure Site-to-Site VPN: Established VPN connectivity to securely monitor Azure VMs from Prometheus servers hosted on AWS.
AWS Production Account: Prometheus servers deployed in AWS production account for high availability and scalability.
Key Responsibilities:
Terraform for Prometheus Deployment: Developed Terraform scripts to automatically deploy and configure Prometheus servers on AWS, ensuring a scalable monitoring solution.
Automated Telegraf Deployment with Ansible: Created Ansible roles to deploy Telegraf on new EC2 instances and Azure VMs automatically during instance provisioning via Terraform, ensuring that all instances were integrated into the monitoring system immediately.
Monitoring Configuration: Configured Prometheus to scrape metrics from Telegraf agents across both AWS and Azure environments and set up alert rules for critical system metrics.
Grafana Dashboard Setup: Built detailed Grafana dashboards for visualizing metrics across both cloud environments, providing insights into performance, resource usage, and system health.
Alerting and Incident Management: Set up Alertmanager to route alerts to PagerDuty, ensuring timely incident responses through Microsoft Teams notifications.
VPN Management: Utilized the existing site-to-site VPN to securely connect the Azure VMs to Prometheus servers in AWS for cross-cloud monitoring.
Continuous Integration with Terraform and Ansible: Integrated Terraform and Ansible to ensure infrastructure provisioning and monitoring setup were fully automated, minimizing manual intervention.
Project Outcomes
Full Automation of Monitoring Deployment: Successfully automated the deployment of Prometheus servers using Terraform and the deployment of Telegraf agents using Ansible during instance provisioning, ensuring seamless and consistent monitoring across AWS and Azure.
Efficient Monitoring Across Clouds: Deployed a centralized monitoring system that provided real-time data on instances across both AWS and Azure, enhancing visibility into system performance.
Real-Time Incident Response: Implemented a robust alerting system using Prometheus, Alertmanager, and PagerDuty, improving response times through integrated notifications on Microsoft Teams.
Scalable and Secure Monitoring Infrastructure: Deployed scalable Prometheus servers in AWS production and ensured secure monitoring of Azure VMs via site-to-site VPN.
Proactive Issue Detection: Enabled proactive monitoring and issue detection with custom Grafana dashboards, allowing teams to identify and resolve issues before they escalated.
Latest Projects
Things I do

CLOUD
Amazon Web Services
AZURE Cloud
Google Cloud Platform
IBM Cloud
DevOps Tools
Jenkins
Ansible for software provisioning
Packer for automated machine images
AWS CodePipeline
AWS CodeDeploy
Azure Pipelines
Chef
Github Actions
GitLab CI/CD
CircleCI
Microservices
Kubernetes
Docker
AWS ECS
AWS EKS
AKS
Azure Container Apps
Monitoring and Alerting
Telegraf
Prometheus
Grafana
Pagerduty
Infrastructure as Code
CloudFormation
Terraform
AWS CDK (Python & Typescript)
Operating System
Most of Linux Server Distributions
Window Server 2019/2022
Databases
MySQL
MariaDB
MongoDB
PostgreSQL
MSSQL
Scripting
Shell Scripting
Python
YAML
PowerShell
My Certifications


Let's Talk
Book a call to talk regarding the current vacancy at your organization.
@2024 RAJ PATEL