What DevOps is
For most of the history of software, developers were measured on how much they shipped and operations on how little broke. Those goals pull in opposite directions, so changes queued up, releases became large and risky, and every outage started with an argument about whose fault it was.
DevOps removes that wall. One team owns a service end to end and is measured on outcomes both sides care about: how quickly a change reaches users and how reliably the service runs. To make that possible, teams automate the path to production, release in small batches, watch production closely, and learn from every failure.
It is first a change in how people work together, then a set of practices, and only then a set of tools.
What DevOps is not
Principles
A useful checklist is CALMS. If one of these is missing, the others stall.
The DevOps Handbook frames the same ideas as three ways: optimize the flow of work from development to production, create fast feedback from production back to development, and build a culture of continual experimentation and learning.
How to measure it
Research from the DORA program, published in the book Accelerate, found four measures that predict software delivery performance. Two measure speed and two measure stability, and the best teams improve both together.
Measure per team and per service, and watch trends over months. Comparing teams against each other, or tying the numbers to bonuses, gets you gamed metrics instead of better software.
Core practices
Bringing DevOps into your organization
DevOps is adopted one team and one bottleneck at a time. The phases below are a typical path. Timings vary with the size of the organization, but the order matters: skipping the foundations to build a platform nobody is ready to use is the most expensive mistake on this page.
- 1AssessWeeks 1–4
Understand where time goes today and pick one place to start.
- Map the value stream for one service: every step from an idea to running in production, with how long work waits at each step.
- Record a baseline of the four metrics above, even if you measure them roughly by hand.
- Choose one pilot team that owns a real, customer-facing service and wants to change. Avoid a greenfield toy project.
- Agree with leadership on the outcome you are after, in business terms: faster time to market, fewer outages, less overtime.
Done whenYou can show where the waiting happens and have a pilot team with time set aside for the work. - 2FoundationsFirst 90 days
Make changes small, safe and visible for the pilot team.
- Put everything in version control: application, infrastructure, configuration, pipeline definitions and runbooks.
- Run CI on every pull request. Keep branches short-lived and merge to the main branch at least daily.
- Automate deployment to at least one environment, triggered from the pipeline, never from a laptop.
- Add basic monitoring and alerting on what users feel: errors, latency, availability.
- Hold a blameless postmortem for every significant incident and track the follow-up actions to completion.
Done whenThe pilot team deploys on demand without a release meeting, and knows within minutes when production breaks. - 3ScaleMonths 3–12
Make the practices repeatable and spread them to more teams.
- Provision environments with infrastructure as code. Treat a manual change in production as an incident.
- Deploy to production through the pipeline with progressive delivery: canary or blue-green releases and automated rollback.
- Define service level objectives (SLOs) for key services and use error budgets to balance features against reliability.
- Share on-call between developers and operators for the services they build together.
- Move security left: scan dependencies, images and infrastructure code in the pipeline, and manage secrets centrally.
- Let the pilot team teach the next teams. Pairing beats training slides.
Done whenSeveral teams deploy independently, and reliability is discussed with numbers from SLOs instead of opinions. - 4PlatformYear 2 onwards
Reduce the effort every team spends on the same plumbing.
- Form a platform team that builds paved roads: a templated service, pipeline, observability and deployment that work out of the box.
- Offer the platform as a product with documentation, support and a roadmap driven by its users, the development teams.
- Keep the paved road optional but attractive, so teams choose it because it is easier.
- Measure developer experience alongside the four metrics: onboarding time, time to first deploy, how often people wait on others.
Done whenA new service goes from idea to production in a day, and teams spend their time on product work, not plumbing.
How to organize teams
Structure follows the work. The model described in the book Team Topologies fits DevOps well: a few team types with clear responsibilities, and as few handoffs between them as possible.
The anti-pattern to avoid is a separate DevOps team that sits between development and operations and takes tickets from both. It recreates the wall DevOps is meant to remove.
DevOps and SRE
Site Reliability Engineering (SRE) is a concrete way to run DevOps, which started at Google, where operations work is treated as a software problem. Its core tools work in any organization:
Common mistakes
Start on Monday
You don’t need a transformation program to begin. Five things one team can do this week:
- Pick one service and write down every step from commit to production, with how long each one takes.
- Measure how often that service deploys and how long a change takes to reach users.
- Find the longest wait in that flow and fix only that.
- Run your next incident review without blame, and ask “what made this possible?” instead of “who did this?”.
- Book a recurring slot for improvement work and protect it.