Senior Kubernetes and Open Source Platform Engineer
India
Full Time
Experienced
Senior Kubernetes and Open Source Platform Engineer
Zappsec Technologies Inc.
About the role
Zappsec advises clients who run their own Kubernetes platforms. Their engineers operate the clusters. You tell them how to operate them well, and you take the problems they cannot solve.
The platforms in scope are self managed rather than cloud managed. Expect Rancher RKE2 on bare metal or private infrastructure, with Harbor, GitLab Runner, MetalLB, and OpenSearch around it.
The work splits three ways. You answer escalations, you review platform health and upgrade plans, and you teach the client team what you find. You will rarely hold production admin rights.
What you will do
Advisory and escalation support
Zappsec Technologies Inc.
About the role
Zappsec advises clients who run their own Kubernetes platforms. Their engineers operate the clusters. You tell them how to operate them well, and you take the problems they cannot solve.
The platforms in scope are self managed rather than cloud managed. Expect Rancher RKE2 on bare metal or private infrastructure, with Harbor, GitLab Runner, MetalLB, and OpenSearch around it.
The work splits three ways. You answer escalations, you review platform health and upgrade plans, and you teach the client team what you find. You will rarely hold production admin rights.
What you will do
Advisory and escalation support
- Take escalations on Kubernetes, RKE2, Harbor, GitLab Runner, MetalLB, and OpenSearch.
- Diagnose problems from logs, metrics, manifests, and screen shares rather than direct console access.
- Work alongside client engineers while they apply the fix, and explain why it works.
- Reproduce complex failures in a lab environment when the client environment cannot be disturbed.
- Decide when an issue belongs upstream, then prepare the bug report or vendor case with the evidence attached.
- Track recurring escalations and name the underlying platform weakness that produces them.
- Diagnose control plane issues covering etcd health, API server latency, scheduler behaviour, and certificate expiry.
- Advise on cluster topology, node pool design, and failure domain placement.
- Review RBAC, admission control, Pod Security Admission, and network policy design.
- Troubleshoot CNI and CSI behaviour, including pod networking, MTU problems, and volume attachment failures.
- Guide etcd backup strategy and test restore procedures against a real recovery target.
- Advise on CIS hardening profiles and air gapped installation patterns where the client requires them.
- Harbor. Review project structure, robot accounts, replication rules, scanning policy, garbage collection, and storage growth.
- GitLab Runner. Advise on the Kubernetes executor, concurrency limits, caching, resource requests, and build isolation choices.
- MetalLB. Review layer 2 and BGP mode selection, address pool design, and peering with upstream routers.
- OpenSearch. Advise on shard and index strategy, index lifecycle policy, snapshot repositories, node roles, and cluster sizing.
- Flag where a component has been pushed past its design intent, and say what should replace it.
- Produce upgrade plans covering version skew, sequencing, prerequisites, and rollback.
- Detect deprecated and removed API usage before an upgrade rather than during it.
- Review Helm chart and CRD upgrade paths, including the ones that cannot be rolled back.
- Run periodic platform health reviews covering capacity, resource limits, backup validity, certificate lifetimes, and security posture.
- Write each review as a prioritized finding list with a remediation step per finding.
- Give the client an honest severity call, including the findings they will not want to hear.
- Run periodic advisory sessions for the client platform team on a set cadence.
- Teach from real incidents in the client environment rather than generic material.
- Leave written runbooks and decision records behind, not slide decks alone.
- Build the client team's capability so escalation volume falls over the engagement.
- Six or more years in platform or infrastructure engineering, with three or more running Kubernetes in production.
- Direct operational experience with self managed Kubernetes on bare metal or private infrastructure.
- Hands-on experience with Rancher or RKE2, including cluster lifecycle and upgrades.
- Working experience with at least three of Harbor, GitLab Runner, MetalLB, OpenSearch, and equivalent open source platform components.
- Proven record of Kubernetes upgrade planning and execution across multiple versions.
- Strong Linux, networking, and storage fundamentals. You can read a packet capture and an etcd metric.
- Ability to diagnose a problem you cannot touch, using evidence the client provides.
- Written communication strong enough that findings reach a client without an editor.
- CKA, CKS, or CKAD certification.
- Experience in air gapped or regulated environments.
- Consulting, professional services, or vendor escalation background.
- GitOps experience with Argo CD or Flux.
- Observability depth across Prometheus, Grafana, and log pipeline design.
- Terraform, Ansible, or Helm authoring experience.
- Upstream contribution to any project in the scope above.
| Category | Tools |
| Orchestration | Kubernetes, Rancher, RKE2, containerd |
| Registry | Harbor, Trivy |
| CI | GitLab, GitLab Runner |
| Networking | MetalLB, Calico or Cilium, CoreDNS, ingress controllers |
| Search and logging | OpenSearch, OpenSearch Dashboards |
| Storage | CSI drivers, Longhorn or equivalent block and object storage |
| Observability | Prometheus, Grafana, Alertmanager |
| Automation | Helm, Terraform, Ansible, Git |
| Delivery | [Confirm ticketing and documentation stack] |
Apply for this position
Required*