DevOps

Deploys on demand. On-call that lets you sleep.

Senior platform engineers for CI/CD, infrastructure-as-code, and observability - so releases stop being events and on-call stops being a tax on your best people.

The brief we usually get

Two problems we are called in to fix.

Concrete starting points, each with the team we field, the stack, and what lands in the first 30 days.

01

Reliability & observability

Typical problemManual deploys, thin monitoring, and an on-call rotation everyone dreads.
What we shipGitOps deploys, SLOs with error budgets, and alerting that pages on symptoms, not noise.
Example stackKubernetesArgoCDPrometheusGrafana
First 30 daysOn-demand deploys plus baseline SLOs and alerting in place.
02

Internal platform & golden paths

Typical problemEvery team wires infra by hand; nothing is repeatable and onboarding takes weeks.
What we shipPaved paths - templates, CI, and self-serve environments behind one workflow.
Example stackTerraformGitHub ActionsBackstageHelm
First 30 daysOne golden path a team can ship a service on end-to-end.
Deliverables

What you get.

Operational outputs that make releases, incidents, and reliability tradeoffs easier to manage.

01 / CommitClear release workflow
02 / GuardSLOs and observability
03 / RecoverRunbooks your team owns
Release

Deployment workflow

A clear path from commit to production with approvals, rollback expectations, and ownership.

Why it mattersReleases become routine instead of special events.
Observe

Observability baseline

Dashboards, alerts, and service health signals mapped to what users actually experience.

Why it mattersOn-call responds to real symptoms, not noisy infrastructure trivia.
Target

Reliability targets

Availability, latency, and incident response expectations agreed with engineering and leadership.

Why it mattersTeams can make tradeoffs with a shared definition of reliability.
Handover

Runbooks and handover

Operational notes, escalation paths, and recovery steps your team can use during incidents.

Why it mattersThe platform remains understandable after the engagement ends.
Stack we work in

The reliability toolchain.

Orchestration

KubernetesHelmNomadECS

CI/CD

GitHub ActionsArgoCDGitLab CIFlux

IaC

TerraformPulumiAnsibleCrossplane

Observability

PrometheusGrafanaLokiOpenTelemetry

Incident

PagerDutySentryRunbooks

Platform

BackstageVaultCloud IAM
How we work

From release pain to reliable operations.

The sequence turns fragile deploys and noisy alerts into a repeatable operating model.

PHASE 01

Assess

Map deploy flow, incident history, and where the error budget actually goes.

Time1-2 weeks
DeliverableReliability assessment
PHASE 02

Pipeline & IaC

GitOps deploys and infrastructure-as-code, with a safe rollback path.

Time2-3 weeks
DeliverableOn-demand deploys
PHASE 03

SLOs & alerting

Define SLOs, wire dashboards, and cut alerts down to symptoms that matter.

TimeOngoing
DeliverableSLOs + clean paging
PHASE 04

Operate & hand off

Runbooks, on-call rotation, and platform docs. Your team owns it.

TimeMonth 4+
DeliverableOwnership handover
Who delivers this

Named engineers from the bench

A snapshot. You interview the actual people before anyone joins your team.

BesnikPlatform engineer
Available now

Cloud / DevOps Engineer

10 yrs · platform & reliability
AWSKubernetesTerraformArgoCD
EnglishC1
GermanB2
InterviewsReady to interview
ArbenReliability engineer
Available now

Site Reliability Engineer

8 yrs · SLOs & incident response
PrometheusGrafanaGoHelm
EnglishC1
GermanB1
InterviewsCleared · 1 client
Request a shortlist

Send us the platform role. We'll return a named shortlist.

Tell us the environment and the pain. Within 5 business days you get 2-4 named engineers with CVs and rates from 290 - 380€/day, to interview yourself.