Skip to content
View Krish9130's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report Krish9130

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Krish9130/README.md

Header

Line 1 Line 2 Line 3 Line 4

LinkedIn Medium YouTube Docker Hub HackerRank Instagram Stack Overflow


Profile Views Followers


About Me

name: Krushna Sonawane
role: AIOps Engineer | Platform Engineer | SRE
location: Pune, India 🇮🇳
experience: DevOps & Cloud Infrastructure
focus:
  - AI-Driven Operations & AIOps
  - Cloud-Native Infrastructure on AWS
  - CI/CD Pipeline Design & GitOps
  - Container Orchestration (Kubernetes/EKS)
  - Infrastructure as Code (Terraform/Ansible)
  - Full-Stack Observability & Monitoring
  - SRE Practices (SLOs, Incident Response)
  - AI Pipeline Deployment & Observability
  - DevSecOps & Security Automation
  - Workflow Automation (n8n, LLM APIs)
currently_learning:
  - Kagent (Kubernetes AI Agents)
  - LangSmith (AI Pipeline Observability)
  - GPU Cluster Operations
  - LLM Gateway Management
  - Platform Engineering
fun_fact: "I automate things before my coffee gets cold ☕"

AIOps Engineer with a strong focus on Site Reliability Engineering (SRE), building scalable and resilient cloud infrastructure on AWS. Proficient in CI/CD pipeline development, containerization using Docker and Kubernetes, and infrastructure automation with Terraform and Ansible. Experienced in implementing observability solutions using Prometheus, Grafana, Datadog, and AWS CloudWatch to improve system availability, reduce MTTR, and optimize performance. Skilled in automation using Python and Bash to reduce operational toil, with growing expertise in AI-driven workflows using n8n, LangSmith, Kagent, and modern AI tools.


📋 Table of Contents

About Skills Experience Projects AIOps Writing Certs Stats Connect


🛠 Tech Stack & Tools

Tech Stack Header

☁️ Cloud & Infrastructure

AWS EC2 EKS ECS S3 RDS Lambda Route 53 CloudWatch CloudFront VPC IAM SNS API Gateway ElastiCache

🐳 Containers & Orchestration

Docker Podman Kubernetes Helm Kustomize Karpenter

⚙️ CI/CD & GitOps

Jenkins AWS CodePipeline GitLab CI GitHub Actions Git GitHub GitLab ArgoCD Argo Rollouts Harness

🏗️ Infrastructure as Code

Terraform CloudFormation Ansible Packer

📈 Monitoring & Observability

Prometheus Grafana Datadog Loki ELK Stack PagerDuty Uptime Kuma

🔒 DevSecOps & Security

SonarQube Trivy Snyk OWASP

🤖 AIOps & AI Automation

n8n LangSmith Kagent ChatGPT Claude Gemini FastAPI

💻 Scripting, Databases & Tools

Python Bash Groovy Linux Nginx Apache Tomcat PostgreSQL MySQL MongoDB Redis RabbitMQ Cockpit Garden YAML Flask


💼 Professional Highlights

Experience Header

Quantified impact from real-world production environments — no company names disclosed.

🏗️ Infrastructure & Cloud

  • 🔻 30% monthly cost reduction — Rebuilt entire infrastructure from scratch with efficient resource provisioning
  • 70% faster environment setup — Provisioned scalable AWS infra via Terraform + Ansible, maintaining 99.9% uptime
  • 📊 Created AWS cost estimation reports for C-level leadership, accelerating project approval cycles
  • 🔧 Authored 800+ lines of Terraform modules covering EKS, RDS, S3, ALB/NLB, Route 53, WAF rules
  • 🏠 Designed multi-account AWS landing zones with Control Tower, Organizations, and SCPs

🚀 CI/CD & Deployment

  • ⏱️ 30% shorter release cycles — Configured and optimized pipelines using Jenkins, GitLab CI, and AWS CodePipeline
  • 🎯 90% fewer deployment errors — Helm chart-based deployments for 30+ microservices via ArgoCD GitOps
  • 🔄 Implemented ArgoCD ApplicationSets with canary → staging → production promotion gates
  • 40% fewer post-release defects — Integrated automated testing frameworks into CI/CD workflows
  • 📦 Managed Docker image lifecycle: multi-stage builds, ECR policies, vulnerability scanning with Trivy + Inspector

📈 Observability & SRE

  • 40% MTTR reduction — Set up alert policies and anomaly detection using CloudWatch, Datadog, and Grafana
  • 🔍 MTTD under 3 minutes — Led incident response as primary on-call with documented post-mortems
  • 📊 Deployed full observability stack: Prometheus (Thanos), Grafana (25+ dashboards), Loki, Tempo, OpenTelemetry
  • 🎯 Defined SLOs and error budgets for critical services (99.9% availability, p99 < 200ms latency)
  • 📉 Reduced time-to-root-cause from 45min to 8min using structured logging and CloudWatch Logs Insights
  • 🔄 Migrated legacy monitoring to centralized Datadog — standardizing alerting, logging, and performance tracking

🔒 Security & Compliance

  • 🛡️ Implemented DevSecOps tooling (SCA, SAST, DAST) — early vulnerability detection in pipelines
  • 🔐 Built centralized secrets management with AWS Secrets Manager + ExternalSecrets Operator — rotated 200+ secrets
  • 🌐 Configured AWS WAF + Shield — blocked 2.3M+ malicious requests in the first month
  • ✅ Maintained SOC 2 Type I compliance with AWS Config, GuardDuty, Security Hub, and CloudTrail

🤖 Automation & Toil Reduction

  • 🕐 20+ hours/month saved — Scripted repeatable operations using Python, Bash, and Ansible
  • Provisioning: 2 hours → 10 minutes — Ansible for 100% environment consistency (dev/staging/prod)
  • 🔧 Custom runbook automation: EKS node recycling, RDS snapshot rotation, IAM key rotation — saving 12+ hrs/week
  • 📋 Produced architecture diagrams and CI/CD documentation, improving onboarding speed

🤖 AIOps & AI Infrastructure

AIOps Header

🔭 AI Pipeline Operations

  • Deployed FastAPI-based LLM inference services on EKS with GPU-backed HPA
  • Set up LangSmith self-hosted tracing for agentic pipeline observability
  • Managed GPU node pool operations with utilization profiling per model type
  • Built n8n workflow automations for AI-driven operational tasks

📊 AI Observability

  • Custom Prometheus exporters for LLM metrics (token usage, latency, error rate, cost per run)
  • LangGraph node-level tracing — execution latency, retry rates, error classifications
  • Loggifly structured log streaming for AI services (model ID, prompt/completion tokens)
  • 14 Prometheus alert rules for AI-specific failures (LLM timeouts, RAG degradation, agent loops)

🚀 AI Deployment Patterns

  • Blue-green & canary rollouts for LLM inference via ArgoCD + Argo Rollouts
  • MCP server deployment on EKS with cert-manager TLS and NGINX Ingress
  • HPA tuning for GPU-backed pods with startup probes for model warm-up
  • Uptime Kuma external monitoring for AI API endpoints

🧠 Agentic Infrastructure

  • Kagent — Kubernetes AI agents for PVC pressure, pod restart loops, node diagnostics
  • Copilot Agent — Infra code review with cost estimation (Infracost)
  • Sencho — AI log anomaly detection, reduced alert noise by 70%
  • AI-driven incident triage — LLM timeout, embedding cold-start, vector DB ops

🔗 AI Engineer ↔ AIOps Alignment

What AI Engineers Build What I Deploy, Integrate & Monitor
LangGraph agentic pipeline (Python/FastAPI) EKS deployment, HPA tuning, rolling updates, health probes
RAG backend with ChromaDB vector store Persistent volume provisioning, backup/restore, storage class
MCP server for LLM orchestration Kubernetes manifests, TLS, cert-manager, ingress, DNS
LangSmith tracing pipeline Self-hosted LangSmith deployment + Grafana ops dashboards
Multi-agent orchestration platform Prometheus exporters, alert rules, incident runbooks
LLM API-powered dashboards API gateway, rate limiting, cost tracking, key rotation

🚀 Featured Projects

Projects Header

🛡️ DevSecOps CI/CD Pipeline

Secure Pipeline for 5 Microservices

  • 🔻 80% less manual deployment time
  • 🚀 3x improved release frequency
  • 🛡️ 100% automated security scanning (SAST + DAST)
  • 🔻 70% fewer production vulnerabilities
  • 📊 Prometheus + cAdvisor + Grafana monitoring with 60% MTTD reduction

🏗️ AWS Infrastructure Automation

Reusable IaC for Multi-Environment Deployments

  • 📦 Reusable Terraform modules for consistency & scalability
  • 🌐 VPC, IAM, EC2, ALB, ASG, S3, Route53, CodePipeline
  • ⚙️ Ansible post-provisioning for EC2 configuration
  • 🔄 Seamless integration across dev/staging/prod

User-Friendly CloudWatch Metrics App

  • 🖥️ Streamlit-based UI for fetching AWS metrics
  • 📥 Download metrics in CSV and JSON format
  • ⏱️ Custom time range selection
  • 📊 No AWS Console navigation required

📱 WhatsApp Alert System

Automated Production Server Monitoring

  • 🤖 Selenium-based WhatsApp notification automation
  • ⏰ Crontab-integrated server health monitoring
  • 🔄 Auto-restart on downtime detection
  • 📲 Real-time WhatsApp alerts on recovery

Serverless Error Alerting Pipeline

  • Lambda monitors CloudWatch log groups
  • Triggers SNS on error patterns
  • Automated incident detection
  • Reduces MTTD for production issues

Automated Backup & One-Click Rollback in Pipelines

  • 🗄️ Automated pre-deploy snapshots stored to S3
  • ♻️ One-click rollback to any previous artifact version
  • 🔔 Failure detection triggers instant rollback in pipeline
  • ⏱️ Jenkins + GitHub Actions integrated execution
  • 🚀 Zero-downtime deployments with versioned rollback

n8n Projects Header

📊 n8n — Cost & Hours Report Pipeline

Automated Cost & Hours Reporting via Email

  • 📋 Reads project hours & cost data from Google Sheets
  • 📊 Generates a professional HTML report with breakdowns
  • 📧 Sends formatted email reports to stakeholders automatically
  • ⏰ Scheduled triggers for daily/weekly cost summaries
  • 💰 Tracks resource utilization and budget variance

🚨 n8n — EOL Alert Management System

Service End-of-Life Detection & Alerting

  • 🔍 Reads service inventory from Google Sheets
  • ⏳ Detects services approaching end-of-life (EOL) dates
  • 🚨 Triggers HTML email alerts with EOL timelines & risk levels
  • 📋 Includes upgrade recommendations & migration paths
  • 🔄 Scheduled scans to catch upcoming EOL windows proactively

📰 n8n — AWS Daily News Updater

Automated AWS News Feed & Notifications

  • 📡 Fetches latest AWS service announcements & updates daily
  • 📝 Summarizes key changes relevant to DevOps/SRE teams
  • 📲 Delivers curated news digest via Telegram
  • 🔔 Filters by service categories (EKS, Lambda, IAM, etc.)
  • ⏰ Scheduled daily triggers for morning briefings

🎨 n8n — DevOps Content Automation

AI-Powered Content Creation & Social Posting

  • 🤖 Auto-generates DevOps-themed images using AI (DALL-E/Stable Diffusion)
  • 📝 Creates engaging DevOps content (tips, best practices, tutorials)
  • 📲 Publishes posts with AI-generated images to Telegram channel
  • ⏰ Scheduled automation for consistent content delivery
  • 🎯 Topics: Kubernetes, Docker, CI/CD, SRE, Cloud, AIOps

📝 Technical Writing

"Documentation is a love letter you write to your future self." — Damian Conway

📰 Article 🏷️ Topics
Automating Directory Backups with Version Control in Jenkins Jenkins Bash CI/CD Backup Automation

Medium


🎯 Currently Working On

Currently Learning
🔄 Focus Area 📚 What I'm Exploring
AIOps & AI Infra Kagent, LangSmith, GPU ops, AI incident triage
GitOps ArgoCD, Argo Rollouts — Canary & Blue-Green for AI services
Platform Engineering Internal Developer Platforms, Golden Paths
DevSecOps SAST/DAST pipeline integration, policy-as-code, secrets rotation
AI Automation n8n workflows, LLM API integration (ChatGPT, Claude, Gemini)

🏅 Certifications

Certs Header

🏅 Certification 🏢 Issuer 📌 Status
AWS Solutions Architect – Professional Amazon Web Services 📖 Preparing
AWS DevOps Engineer – Professional Amazon Web Services 📖 Preparing
Certified Kubernetes Administrator (CKA) CNCF 📖 Preparing
HashiCorp Terraform Associate HashiCorp 📖 Preparing

📊 GitHub Analytics

Stats Header

GitHub Stats GitHub Streak Top Languages

📈 Detailed Profile Analytics (Click to expand)
Profile Stats Repos per Language Most Commit Language
Profile Details
🏆 GitHub Trophies (Click to expand)
GitHub Trophies
📊 Contribution Graph (Click to expand)
Contribution Graph

🌐 Connect with Me

Connect



LinkedIn Medium YouTube Docker Hub


HackerRank Instagram Stack Overflow Gmail



Footer

Visitor Count

Pinned Loading

  1. Krish9130 Krish9130 Public

    Config files for my GitHub profile.

    1

  2. PDF-JPG-Converter PDF-JPG-Converter Public

    JavaScript

  3. Tektoncatalog Tektoncatalog Public

    Forked from tektoncd/catalog

    Catalog of shared Tasks and Pipelines.

    Shell