SYSTEM STATUS: OPERATIONAL

Thejas Kumar Patel

Infrastructure & Site Reliability Engineer

I keep banking-scale systems observable, stable, and self-healing — turning noisy alerts into clean dashboards, and incidents into root causes fixed for good. 15+ years across production support, cloud infrastructure, and enterprise middleware.

15+ yrs in production & infra ops
80% faster incident response via automation
99%+ SLA / KPI compliance maintained
150+ production applications supported
Monitoring & Observability
Automation & Self-Healing
Incident Response
Cloud & Middleware

01 / About

Fifteen years of keeping other people's systems up.

I'm a production support and infrastructure engineer who has spent the last decade and a half inside banking, fintech, and airline environments — the kind where downtime shows up on a balance sheet or a departures board. My work sits at the intersection of site reliability, observability, automation, and incident response: building the dashboards and alerts that catch problems early, and the self-healing scripts that fix the recurring ones before anyone gets paged.

I've supported everything from WebSphere MQ messaging backbones to Azure Red Hat OpenShift clusters, and I've learned that most outages aren't mysteries — they're missing alerts, undocumented dependencies, or manual steps nobody automated yet. Closing those gaps is the job.

Consistently recognized by clients for technical ownership and being a dependable, go-to member of the team.

02 / Deployment Log

Experience

Production Support / Operations Engineer

  • Own production support for 150+ business-critical banking applications, ensuring high availability and consistent SLA compliance.
  • Built and maintained Splunk dashboards for real-time visibility across application, infrastructure, and operational health.
  • Set up proactive alerting across Splunk, AppDynamics, Grafana, Prometheus, and Splunk On-Call, reducing production alert noise by 25–40% through intelligent tuning.
  • Built Power Automate workflows to alert the team the moment an incident landed in our queue, cutting response time by roughly 80%.
  • Owned RCA end-to-end and automated repetitive operational tasks into self-healing, auto-triage workflows — cutting MTTR by 80% and manual triage effort by half, since triage could be decided directly from logs already surfaced rather than dug up by hand.
  • Designed GitHub Copilot–based operational agents that automate morning ops summaries, overnight production reports, incident tracking, and follow-up reminders — saving 1–2 hours of manual effort daily.
  • Piloted Devin AI to independently review production release pull requests and validate configuration changes ahead of deployment.
  • Consistently maintained >99% SLA/KPI compliance across incidents.
  • Kept Autosys-scheduled batch workloads running reliably in coordination with application teams.
SplunkAppDynamicsGrafanaPrometheusSplunk On-CallPower AutomateAutosysGitHub CopilotDevin AI

Cloud / Middleware / Infrastructure Support Engineer

  • Supported BAU deployments and network connectivity across on-prem data centers, Azure Red Hat OpenShift (ARO), and Pivotal Cloud Foundry.
  • Provided on-call incident support focused on network, middleware, and cloud connectivity, coordinating resolution across teams.
  • Integrated Wiz CSPM into ARO clusters using Terraform, improving visibility into exposure and misconfigurations.
  • Provisioned Azure VMs, Cosmos DB, and ARO clusters in new landing zones.
AzureAzure Red Hat OpenShiftTerraformWiz CSPMPCFCosmos DB

Infrastructure Support Engineer

  • Delivered end-to-end environment and application support for the lower environments of a major bank.
  • Stood up monitoring, alerting, and dashboards across on-prem and cloud using Splunk, Dynatrace, and Grafana.
  • Triaged issues spanning network, MQ, HAProxy, API gateway, WAS, Tomcat, DataPower, CyberArk, ADFS, AWS, and Kubernetes.
  • Protected testing windows by ensuring environment availability after every CI/CD deployment.
  • Applied self-healing automation and coordinated response on standard and high-severity incidents.
SplunkDynatraceKibanaGrafanaWASKubernetesAWS

WebSphere MQ & WAS Administrator

  • Designed and implemented WebSphere MQ messaging solutions, including capacity tuning and security for data in transit and at rest.
  • Built high availability for MQ using OS-level and MQ clustering across Linux, Solaris, AIX, HP-UX, and OpenVMS.
  • Managed 1,500+ WebSphere MQ instances supporting cross-platform communication.
  • Handled severity incidents and RCA; ran build and deployment for Java and .NET apps on WebSphere Application Server.
WebSphere MQWASVCS

WebSphere MQ & WAS Administrator

  • Delivered 24/7 support for MQ and WAS infrastructure for the largest global airline alliance.
  • Migrated MQ 6.0 → 7.0 and WAS v6.1 → v7.0 with minimal service disruption.
  • Supported implementation planning, release execution, and post-release stabilization.
WebSphere MQWAS

03 / Stack

Tools & Technologies

Site Reliability & Operations

Production support, incident management, problem management, root cause analysis (RCA), high availability, change management, release coordination, SLA/KPI management

Monitoring & Observability

Splunk, Splunk On-Call, AppDynamics, Grafana, Prometheus, Dynatrace, Kibana, PagerDuty, IR360

Cloud & Platform Engineering

Azure, Azure Red Hat OpenShift, Pivotal Cloud Foundry, AWS, Kubernetes, Terraform, Azure Cosmos DB

Automation & DevOps

Autosys, Power Automate, self-healing & auto-triage scripting, CI/CD environment support, Python, Perl

Middleware & Messaging

WebSphere MQ, WebSphere Application Server, IHS, Tomcat, Apache, WMB, API Gateway, F5, HAProxy, DataPower

Security

Wiz CSPM, CrowdStrike, CyberArk, ADFS

Databases

Oracle, Microsoft SQL Server, MongoDB, DB2

Languages & OS

C, C++, Python, Perl, SQL, HTML/XML · RHEL, Solaris, AIX, OpenVMS, Windows Server

04 / Certifications

Certifications

ISC2 Certified in Cybersecurity
AWS Certified Cloud Practitioner
IBM MQ V7 Administrator
IBM WAS V7 ND Administrator
ITIL Foundation
ManageEngine Log360 Associate

05 / Education

Background

B.Tech

06 / Contact

Let's talk about keeping things running.

Based in Jersey City, NJ. Open to conversations about production support, SRE, and infrastructure roles in banking, fintech, and beyond.