Ethan
Kane
Senior Platform Engineer FanDuel
I provision, bootstrap, and operate EKS clusters at fleet scale. Terraform underneath, ArgoCD on top, and enough automation that standing up a new cluster is a pull request.
$
Live · type help
- Fleet
- 150+
- EKS clusters
- Footprint
- Multi-region
- Regions, local zones, Outposts
- Kubernetes since
- 2020
- Across three companies
- Open source
- 4
- Published projects
About
- Origin
- Came into tech through an unconventional route: Earth Science degree, civil engineering on water infrastructure, then a postgrad in software development. That shapes how I think. I care about what actually breaks in production, not what looks elegant in a design doc.
- Based
-
Edinburgh, Scotland
Originally from Galway, Ireland
- Current
-
Senior Platform Engineer
FanDuel
- Focus
- Kubernetes fleet management, GitOps, cluster automation, and keeping things reliable at scale.
Projects
cellcast
Gocontroller-runtimeOIDCTokenRequestHelmPrometheus
A multi-cluster placement oracle and short-lived credential broker. A deploy pipeline asks which cluster a workload goes to and what credential gets it there; cellcast answers both and stores neither, authenticating the caller with the OIDC token its CI platform already issues. Candidate cells are filtered by policy before they are scored, the credential is minted through the Kubernetes TokenRequest API and expires in minutes, and the whole thing attaches to an existing pipeline instead of replacing it.
OOM Oracle
GoeBPFCO-REcgroupsclient-goHelm
A Kubernetes node agent that explains OOM kills at the level the control plane cannot. Kubernetes reports OOMKilled and exit code 137, then stops. This attaches an eBPF kprobe to oom_kill_process and samples cgroup memory continuously, so a report names which process died, what it held at the moment the kernel chose it, and the memory curve on the way there. Falls back to cgroup polling where BTF is absent, and ships a cosign-signed image and OCI Helm chart.
Scale Sentry
GoKubebuildercontroller-runtimeHelmPrometheus
A Kubernetes operator that validates autoscaling behaviour under load. Drives HTTP, HTTP/2, or gRPC traffic at a target Deployment, measures HPA scale-up latency against an SLA, and correlates request errors with EndpointSlice churn to catch cold-start traffic leakage. Ships as a signed OCI Helm chart with Prometheus metrics and a Grafana dashboard.
kubectl-fleet
Goclient-goCobraGoReleaser
A kubectl plugin for multi-cluster operations. Fans queries out across every kubeconfig context in parallel and merges the results into one context-aware table: incident triage, rollout parity checks, and fleet-wide audits without juggling terminal tabs.
EKS Fleet Architecture
TerraformAWS EKSOutpostsLocal ZonesGolang
Architected the EKS estate to run across AWS regions, local zones, and Outposts. Second-highest contributor to the custom Terraform module that provisions every cluster in the fleet.
ArgoCD Platform Bootstrap
ArgoCDKarpenterKyvernoHelmApp of Apps
Project lead on the ArgoCD implementation. Designed the App of Apps pattern so every new EKS cluster auto-provisions Karpenter node pools, target group bindings, ingress, Kyverno policies, and the full observability stack on day one.
Karpenter Fleet Migration
KarpenterCluster AutoscalerEKSTerraform
Migrated the entire EKS fleet from Cluster Autoscaler to Karpenter, configuring node provisioners per workload class. Outposts clusters excluded due to hardware constraints.
Helmfile to Argo Migration
ArgoCDHelmfileHelmGitOps
Migrated all platform services from Helmfile-based deployments to ArgoCD-managed Applications, establishing GitOps as the single deployment path across the estate.
Greenfield K8s Platform
TerraformTerragruntJenkinsAWS EKSIstioFlagger
Built a complete Kubernetes-native platform from scratch with one other engineer. Ground-up infrastructure: CI/CD pipelines, networking, service mesh, and progressive delivery, the full stack from zero.
Experience
2022 - Present
Current
Senior Platform Engineer
FanDuel
- 01 Platform Engineering and Compute Operations teams: provisioning and managing a fleet of 150+ EKS clusters across AWS regions and Outposts
- 02 Designed and implemented the ArgoCD App of Apps pattern to bootstrap all core cluster utilities: ingress controllers, cert-manager, external-dns, cluster autoscaling, and the Datadog Operator
- 03 Every cluster arrives with a full set of Datadog monitors and dashboards on day one, driven by operator CRs in a centralised utilities repo
- 04 Deployed and tuned Karpenter across the fleet, replacing Cluster Autoscaler with node provisioner configs tailored per workload class
- 05 Built Bottlerocket bootstrap container pipelines for custom node initialisation, enabling secure and reproducible cluster join flows
- 06 Wrote Golang tooling for internal platform automation: cluster provisioning workflows, fleet reconciliation, and operational utilities
- 07 Maintained and extended CI/CD pipelines for platform components, driving consistent delivery across all environments
2021 - 2022
DevOps Engineer
Infostretch / Saggezza
- 01 Placed on a US client engagement as the sole DevOps engineer within a small cross-functional team
- 02 Built all CI/CD pipelines from scratch, provisioning AWS infrastructure and Kubernetes workloads via IaC
- 03 Implemented and owned the full AWS environment (EKS, networking, IAM, and supporting services), including a Prometheus, Grafana, and Loki monitoring stack
- 04 Started as an associate developer working Java backend before moving into the platform role
2020 - 2021
Technical Support Engineer
10x Future Technologies
- 01 Supported a cloud-native AWS banking platform built on microservices running on Kubernetes
- 02 Triaged and managed developer tickets across PROD, INT, E2E, and NFT environments
- 03 Led high-severity incident response: communication, root cause analysis, and post-mortems
2019 - 2020
Junior Civil Engineer
Irish Water
- 01 Worked on a watermain rehabilitation project in Galway City
- 02 Transitioned into software development by pursuing a postgraduate degree concurrently
Skills
Stack
01 Kubernetes
02 AWS
03 EKS
04 Karpenter
05 Terraform
06 Terragrunt
07 Helm
08 ArgoCD
09 Golang
10 Python
11 Bash
12 Java
13 Docker
14 Linux
15 GitHub Actions
16 Jenkins
Operating principles
- 01 Operational excellence: zero downtime, zero surprises
- 02 Simpler is better, especially at scale
- 03 Risk-first thinking before every change
- 04 Calm troubleshooting under pressure
- 05 Wide-scale platform impact over local optimisation
Education
2020
Galway, Ireland
Higher Diploma in Software Design and Development (NFQ Level 8)
National University of Ireland, Galway
Final project: mobile study application built with React Native
Grade 2.1
2018
Edinburgh, UK
Postgraduate Diploma in Climate Change: Impacts and Mitigation (NFQ Level 9)
Heriot-Watt University
Dissertation: review of microalgae as sustainable feedstock for biofuel production
2015
Galway, Ireland
BSc Earth and Ocean Science (NFQ Level 8)
National University of Ireland, Galway
Dissertation: distribution of Radium 223/224 as tracers of groundwater mixing in Galway Bay
Grade 2.1
Contact
Open to
Interesting platform engineering conversations, consulting, or just talking Kubernetes.