Cloud · Platform · Applied AI
Andy Barber
20+ years in infrastructure, a decade on Azure. I design and operate reliable Kubernetes platforms and ship applied AI into production — most recently cutting Azure spend 30% and automating away out-of-hours database callouts.
About
Infrastructure, reliability and applied AI
I’m a senior platform and reliability engineer with 20+ years in infrastructure and more than a decade on Microsoft Azure — now building applied AI on the same platform.
My work sits at the intersection of infrastructure, reliability and applied AI: AKS, Terraform, observability and SRE practice, alongside production AI systems — retrieval-augmented generation, computer vision and document intelligence.
I care about systems that are boring in the best sense — observable, reproducible and easy to hand to the next person — and about replacing risky manual processes with self-service automation.
Skills
What I work with
The platforms, tools and practices I rely on to keep production systems reliable.
Cloud
Platform / Containers
Infrastructure as Code
Reliability / Operations
Development
Applied AI
Experience
Career highlights
A few of the teams and problems I’ve worked on — outcomes over job descriptions.
Senior Site Reliability Engineer
Sorted Group
- Built a Slack-driven self-service SQL change platform (PowerShell Durable Functions, durable timers, automated recovery and firewall cleanup) that removed out-of-hours database callouts entirely.
- Authored reusable Terraform modules adopted as the standard Azure deployment pattern across 12 subscriptions.
- Embedded DevSecOps scanning and blue/green + canary releases; built the FinOps practice behind a 30% reduction in Azure spend.
Cloud Solution Engineer
Extrinsica Global
- Designed a production Azure computer-vision / OCR pipeline processing several hundred thousand documents for PII classification, governed with Purview.
- Built a retrieval-augmented documentation assistant as a Microsoft Word add-in — RAG applied to a real authoring workflow.
- Architected a self-service image pipeline that cut customer AVD image builds from weeks to hours (Packer, Docker, Vue.js + C#, Azure Compute Gallery).
Implementation / Infrastructure & SRE Engineer
CluedIn
- Deployed the CluedIn MDM platform on AKS for enterprise customers including Bayer, Microsoft and Svevia.
- Led KEDA event-driven autoscaling with custom Helm charts, cutting thousands per month from Kubernetes compute spend.
- Standardised design and deployment artefacts (Helm, ARM/Bicep) improving installation consistency and supportability.
Support Engineer, Azure Monitor (EMEA)
Microsoft
- Supported enterprise customers implementing Azure Monitor, Application Insights and Log Analytics.
- Analysed telemetry and logs to isolate performance bottlenecks and improve availability for large-scale platforms.
Customer Success Engineer
WANdisco / Cirata
- Contributed to a published migration of 500TB across 8.6M files from a live 800-node Hadoop cluster to AWS S3 with zero business disruption.
- Built Logstash ingest pipelines and Kibana dashboards correlating tens of thousands of log lines into a repeatable diagnostic capability.
Certifications
Current credentials
Backed by current, verifiable credentials.
Microsoft Certified: Azure Solutions Architect Expert
Microsoft
Microsoft Certified: Azure AI Engineer Associate
Microsoft
HashiCorp Certified: Terraform Associate
HashiCorp
-
Microsoft Certified: Azure Administrator Associate (AZ-104) Microsoft
-
Microsoft Certified: Azure AI Fundamentals (AI-900) Microsoft
-
Microsoft Certified: Azure Fundamentals (AZ-900) Microsoft
-
Microsoft Certified: Azure Data Fundamentals (DP-900) Microsoft
-
GitHub Certified: Foundations · Actions · Administration GitHub
Projects
Selected work
A few things I’ve built, from production platforms to practical AI.
Herb Hub 365
A self-hosted, largely autonomous greenhouse platform. Raspberry Pi sensors and cameras feed an event-driven pipeline of Go microservices that turn telemetry into daily AI-written posts, audio narration and timelapse videos — and drive automated watering.
- Event-driven Go microservices on a RabbitMQ backbone (sensor snapshots, watering, video).
- Local-first AI: Ollama with a Gemini fallback, Kokoro TTS narration, MuseTalk avatar video.
- Terraform-managed Azure hosting, with Prometheus + Grafana observability.
MPWrap
A Neovim plugin with an interactive side panel for working with MicroPython devices via mpremote — filesystem browser, live REPL and one-key device actions.
- Three-pane side panel: action menu, filesystem browser and live MicroPython REPL.
- Non-blocking mpremote operations with upload, download and exact-mirror directory sync.
- Keyboard-first workflow with a `:checkhealth` provider for diagnostics.
Azure Platform & Landing Zones
Design and operation of Azure landing zones, AKS platforms, Infrastructure as Code and production operations.
AI Documentation Assistant
A retrieval-augmented documentation assistant delivered as a Microsoft Word add-in — grounded, cited answers inside the authoring workflow.
Computer Vision / OCR Pipeline
A production Azure AI Services pipeline processing several hundred thousand documents for OCR and PII classification, governed with Microsoft Purview.
Contact
Let’s build something reliable.
I’m open to interesting platform, reliability and applied-AI problems. The fastest way to reach me is below.