Day job

Work

I lead cloud reliability on Azure for a multi-tenant SaaS platform: a five-digit number of customer applications across 5 geographical regions. I'm also a founding engineer of the platform underneath it. These are the areas I've owned, as a lead, an implementer, or both.

  1. Reliability & Disaster Recovery

    Backup availability, incident drills, on-call team and process, runbooks, plus large-scale migrations across a broad product portfolio.

    • DR & backups
    • On-call & process
    • Runbooks
    • Incident drills
    • Capacity planning
  2. Engineering Leadership

    Founded the SRE function from the first hire and grew it to a team of four, sole lead through roughly the first eighteen months. Runs hiring and on-call onboarding, owns the incident process, and convenes the cross-team operating cadence across the cloud group. Process owner for the daily operations of a high-density platform.

    • Founded the SRE team
    • Hiring & onboarding
    • Incident process
    • On-call program
    • Cross-team cadence
    • Daily-ops ownership
  3. Data & Analytics Platform

    Business-insight data pipelines on Databricks (Delta tables, jobs), bootstrapped with Terraform. Canonical data models with a semantic data dictionary so queries are correct by construction, and dashboards treated as verifiable query-to-widget contracts rather than pretty pictures.

    • Databricks
    • Delta tables
    • Terraform
    • Canonical modeling
    • Data dictionary
    • Dashboards-as-contracts
    • Fleet analytics
  4. AI Agent Engineering

    Shipped an SRE diagnostics agent and the read-only troubleshooting API + MCP layer it runs on, so agents investigate platform data safely behind hard guardrails. More broadly, building the harness around AI agents rather than just using them: deterministic guardrails (read-only allowlists, completion gates, human-held deploy steps), multi-agent code-review pipelines, and incident investigation encoded as composable skills.

    • SRE diagnostics agent
    • Troubleshooting API + MCP
    • Agent guardrails
    • Multi-agent review
    • Human-in-the-loop
  5. Monitoring & Observability at scale

    Azure infrastructure monitoring across a large estate. Led the strategy, built the implementation, and set the operating processes.

    • Azure Monitor
    • Lead
    • Implementation
    • Process
  6. Infrastructure, Identity & Deployment

    Infrastructure, networking, and identities defined as code, including service accounts and identity management, plus multi-regional service deployments into separate deployment units, with the pipelines and package/artifact management behind them.

    • Terraform
    • Networking
    • Identity management
    • Multi-region
    • Deployment units
    • CI/CD
    • Artifact management
  7. Platform & API Engineering

    Self-documenting internal data and platform APIs (consistent resource conventions, OpenAPI specs, health endpoints) so consumers never have to read the source. Broad Cloudflare platform use (Workers, R2, D1, KV, Queues) and TLS / managed-hostname lifecycle automation at scale.

    • Internal APIs
    • OpenAPI
    • Cloudflare Workers
    • R2 / D1 / KV
    • Cert-lifecycle automation
  8. Edge & Security

    Migrated the multi-tenant edge to Cloudflare for SaaS: WAF, Bot Management, DDoS protection, and Zero Trust access. Manages TLS for thousands of customer custom hostnames as an explicit certificate lifecycle, plus DNS across 10+ zones and email deliverability and anti-phishing (DMARC).

    • Cloudflare for SaaS
    • WAF / Bot Management
    • Zero Trust
    • TLS for custom hostnames
    • DNS (10+ zones)
    • DMARC / email
  9. Governance, Compliance & FinOps

    Azure Policy for compliance, plus account management and cost reporting across Azure (CSP), Cloudflare, and incident.io, alongside day-to-day platform ops on VMs and Azure SQL / elastic pools.

    • Azure Policy
    • Cost reporting
    • Azure CSP
    • Cloudflare
    • incident.io
    • Azure SQL