Day job
Work
I lead cloud reliability on Azure for a multi-tenant SaaS platform: a five-digit number of customer applications across 5 geographical regions. I'm also a founding engineer of the platform underneath it. These are the areas I've owned, as a lead, an implementer, or both.
-
Reliability & Disaster Recovery
Backup availability, incident drills, on-call team and process, runbooks, plus large-scale migrations across a broad product portfolio.
-
Engineering Leadership
Founded the SRE function from the first hire and grew it to a team of four, sole lead through roughly the first eighteen months. Runs hiring and on-call onboarding, owns the incident process, and convenes the cross-team operating cadence across the cloud group. Process owner for the daily operations of a high-density platform.
-
Data & Analytics Platform
Business-insight data pipelines on Databricks (Delta tables, jobs), bootstrapped with Terraform. Canonical data models with a semantic data dictionary so queries are correct by construction, and dashboards treated as verifiable query-to-widget contracts rather than pretty pictures.
-
AI Agent Engineering
Shipped an SRE diagnostics agent and the read-only troubleshooting API + MCP layer it runs on, so agents investigate platform data safely behind hard guardrails. More broadly, building the harness around AI agents rather than just using them: deterministic guardrails (read-only allowlists, completion gates, human-held deploy steps), multi-agent code-review pipelines, and incident investigation encoded as composable skills.
-
Monitoring & Observability at scale
Azure infrastructure monitoring across a large estate. Led the strategy, built the implementation, and set the operating processes.
-
Infrastructure, Identity & Deployment
Infrastructure, networking, and identities defined as code, including service accounts and identity management, plus multi-regional service deployments into separate deployment units, with the pipelines and package/artifact management behind them.
-
Platform & API Engineering
Self-documenting internal data and platform APIs (consistent resource conventions, OpenAPI specs, health endpoints) so consumers never have to read the source. Broad Cloudflare platform use (Workers, R2, D1, KV, Queues) and TLS / managed-hostname lifecycle automation at scale.
-
Edge & Security
Migrated the multi-tenant edge to Cloudflare for SaaS: WAF, Bot Management, DDoS protection, and Zero Trust access. Manages TLS for thousands of customer custom hostnames as an explicit certificate lifecycle, plus DNS across 10+ zones and email deliverability and anti-phishing (DMARC).
-
Governance, Compliance & FinOps
Azure Policy for compliance, plus account management and cost reporting across Azure (CSP), Cloudflare, and incident.io, alongside day-to-day platform ops on VMs and Azure SQL / elastic pools.