Guide

What is DevOps?

DevOps is a way of building and running software in which the people who write the code and the people who operate it share responsibility for the whole lifecycle, from the first commit to running in production at three in the morning.

This guide covers what DevOps is, how to tell whether it is working, and how to introduce it into an organization step by step, including the mistakes that make most transformations stall.

What DevOps is

For most of the history of software, developers were measured on how much they shipped and operations on how little broke. Those goals pull in opposite directions, so changes queued up, releases became large and risky, and every outage started with an argument about whose fault it was.

DevOps removes that wall. One team owns a service end to end and is measured on outcomes both sides care about: how quickly a change reaches users and how reliably the service runs. To make that possible, teams automate the path to production, release in small batches, watch production closely, and learn from every failure.

It is first a change in how people work together, then a set of practices, and only then a set of tools.

What DevOps is not

A job title or a new teamRenaming the ops team to “DevOps” or adding a DevOps team between developers and operations creates one more handoff. The goal is fewer handoffs.
A tool you can buyTools make good habits cheap to repeat. They do not create the habits. A CI server nobody trusts, or a Kubernetes cluster only one person understands, changes nothing.
Getting rid of operationsOperations knowledge becomes more important, not less. It moves into code, platforms and shared on-call instead of tickets.
Speed at the cost of stabilityTeams that deploy more often also tend to have fewer and shorter outages. Small, frequent, reversible changes are what make both possible.

Principles

A useful checklist is CALMS. If one of these is missing, the others stall.

C
CultureShared ownership of the service, from commit to production. Developers carry the pager for what they build, and incidents are reviewed without blame.
A
AutomationAnything done twice by hand is a candidate: builds, tests, environments, deployments, rollbacks, access requests.
L
LeanSmall batches, limited work in progress, and removing waiting time between steps. Most lead time is waiting, not working.
M
MeasurementDecisions are based on data from the delivery pipeline and from production, not on opinions.
S
SharingRunbooks, dashboards, postmortems and pipeline code are open to every team, so lessons travel.

The DevOps Handbook frames the same ideas as three ways: optimize the flow of work from development to production, create fast feedback from production back to development, and build a culture of continual experimentation and learning.

How to measure it

Research from the DORA program, published in the book Accelerate, found four measures that predict software delivery performance. Two measure speed and two measure stability, and the best teams improve both together.

SpeedDeployment frequencyHow often you ship to production.
SpeedLead time for changesTime from commit to running in production.
StabilityChange failure rateShare of deployments that cause a failure in production.
StabilityTime to restore serviceHow long it takes to recover from a failed deployment or outage.

Measure per team and per service, and watch trends over months. Comparing teams against each other, or tying the numbers to bonuses, gets you gamed metrics instead of better software.

Core practices

Bringing DevOps into your organization

DevOps is adopted one team and one bottleneck at a time. The phases below are a typical path. Timings vary with the size of the organization, but the order matters: skipping the foundations to build a platform nobody is ready to use is the most expensive mistake on this page.

  1. 1
    AssessWeeks 1–4

    Understand where time goes today and pick one place to start.

    • Map the value stream for one service: every step from an idea to running in production, with how long work waits at each step.
    • Record a baseline of the four metrics above, even if you measure them roughly by hand.
    • Choose one pilot team that owns a real, customer-facing service and wants to change. Avoid a greenfield toy project.
    • Agree with leadership on the outcome you are after, in business terms: faster time to market, fewer outages, less overtime.
    Done whenYou can show where the waiting happens and have a pilot team with time set aside for the work.
  2. 2
    FoundationsFirst 90 days

    Make changes small, safe and visible for the pilot team.

    • Put everything in version control: application, infrastructure, configuration, pipeline definitions and runbooks.
    • Run CI on every pull request. Keep branches short-lived and merge to the main branch at least daily.
    • Automate deployment to at least one environment, triggered from the pipeline, never from a laptop.
    • Add basic monitoring and alerting on what users feel: errors, latency, availability.
    • Hold a blameless postmortem for every significant incident and track the follow-up actions to completion.
    Done whenThe pilot team deploys on demand without a release meeting, and knows within minutes when production breaks.
  3. 3
    ScaleMonths 3–12

    Make the practices repeatable and spread them to more teams.

    • Provision environments with infrastructure as code. Treat a manual change in production as an incident.
    • Deploy to production through the pipeline with progressive delivery: canary or blue-green releases and automated rollback.
    • Define service level objectives (SLOs) for key services and use error budgets to balance features against reliability.
    • Share on-call between developers and operators for the services they build together.
    • Move security left: scan dependencies, images and infrastructure code in the pipeline, and manage secrets centrally.
    • Let the pilot team teach the next teams. Pairing beats training slides.
    Done whenSeveral teams deploy independently, and reliability is discussed with numbers from SLOs instead of opinions.
  4. 4
    PlatformYear 2 onwards

    Reduce the effort every team spends on the same plumbing.

    • Form a platform team that builds paved roads: a templated service, pipeline, observability and deployment that work out of the box.
    • Offer the platform as a product with documentation, support and a roadmap driven by its users, the development teams.
    • Keep the paved road optional but attractive, so teams choose it because it is easier.
    • Measure developer experience alongside the four metrics: onboarding time, time to first deploy, how often people wait on others.
    Done whenA new service goes from idea to production in a day, and teams spend their time on product work, not plumbing.

How to organize teams

Structure follows the work. The model described in the book Team Topologies fits DevOps well: a few team types with clear responsibilities, and as few handoffs between them as possible.

Stream-aligned teamsOwn a product or service end to end: build it, ship it, run it. Most of your engineers should be here.
Platform teamBuilds and runs the internal platform that stream-aligned teams use through self-service, not tickets.
Enabling teamExperts in an area such as CI/CD, observability or security, who coach teams for a few weeks and then move on.
SREEngineers who apply software engineering to reliability. They can work embedded in teams, as a central team for critical services, or as consultants.

The anti-pattern to avoid is a separate DevOps team that sits between development and operations and takes tickets from both. It recreates the wall DevOps is meant to remove.

DevOps and SRE

Site Reliability Engineering (SRE) is a concrete way to run DevOps, which started at Google, where operations work is treated as a software problem. Its core tools work in any organization:

Service level objectivesA target for how reliable a service must be, measured from the user’s side, for example 99.9% of requests succeed within 300 ms over 28 days.
Error budgetsThe unreliability the SLO allows. While budget remains, ship features. When it runs out, reliability work comes first. This turns the dev versus ops argument into a shared, numeric rule.
Toil limitsManual, repetitive operational work is capped (Google uses 50% of an engineer’s time) so there is always time to automate it away.
Blameless postmortemsEvery significant incident gets a written review focused on causes in the system, not on people, with tracked follow-up actions.

Common mistakes

Renaming instead of changingThe ops team becomes the “DevOps team”, tickets stay tickets, and nothing about how work flows has changed.
Tools firstBuying a platform before agreeing on the problem it solves. Start with the value stream, then choose tools.
Big-bang transformationReorganizing every team at once. Start with one team, prove results, and spread what worked.
Metrics as targetsWhen deployment frequency becomes a target, teams split deploys to hit the number. Use the metrics to find problems, never to rank teams or people.
Automating a broken processAutomating five approval steps gives you five fast approval steps. Remove the waste first.
No time for improvementIf every sprint is 100% feature work, the toil never shrinks. Protect time for pipeline, reliability and tooling work.
Leaving security and compliance outSecurity brought in at the end blocks releases. Bring them in early and encode their rules in the pipeline.

Start on Monday

You don’t need a transformation program to begin. Five things one team can do this week:

  1. Pick one service and write down every step from commit to production, with how long each one takes.
  2. Measure how often that service deploys and how long a change takes to reach users.
  3. Find the longest wait in that flow and fix only that.
  4. Run your next incident review without blame, and ask “what made this possible?” instead of “who did this?”.
  5. Book a recurring slot for improvement work and protect it.