On 28 February 2017, an engineer at Amazon ran a routine command to remove a few servers from a billing system. One input was typed wrong, far more servers were removed than intended, and a large part of the internet that relied on S3 in the us-east-1 region went down for hours. AWS later described the cause in its public post-event summary and added safeguards so the same command could not remove that much capacity so quickly.

The lesson is not “engineers make typos”. The lesson is that systems must be built so a single mistake cannot become a disaster. That is what DevOps is about, and in 2026 every software engineer is expected to understand the basics, not only the “ops person”.

In this note you will learn:

  • what DevOps really means, in one picture,
  • seven basics (CI/CD, infrastructure as code, secrets, IAM, scaling, monitoring, backups),
  • the AWS service behind each one,
  • real failure scenarios and the code or config that prevents them.

Reading time: about 11 minutes.


What DevOps really means

DevOps is not a job title or a tool. It is a way of working where the people who build software also take part in running it, supported by automation. The goal: ship small changes often, safely, and know quickly when something breaks.

1. Code Git + pull requests 2. Build CodeBuild: lint + tests 3. Package Docker image in ECR 4. Deploy CodeDeploy / ECS rollout 5. Run ECS Fargate or Lambda 6. Observe CloudWatch logs + alarms feedback: what we learn in production shapes the next change Infrastructure as code (CloudFormation / CDK / Terraform) describes all of the above IAM roles + Secrets Manager: every arrow above has narrow permissions
The whole DevOps loop on AWS. The rest of this note walks through it, one box at a time.

One more idea before we start: the shared responsibility model. AWS secures the data centres, hardware and core services. You are responsible for what you build on top: your code, your permissions, your configuration and your data. Almost every story below is a failure on the “you” side.


1. CI/CD: never deploy by hand

Continuous Integration (CI) means every change is automatically built and tested. Continuous Delivery (CD) means passing builds are deployed through the same repeatable steps every time. On AWS: CodePipeline orchestrates, CodeBuild builds and tests, CodeDeploy or ECS rolls out. (GitHub Actions works just as well and can deploy to AWS using an IAM role.)

Failure scenario. In 2012 the trading firm Knight Capital deployed new software by copying it to its servers by hand. By most public accounts it reached seven of eight servers, and old, dormant code on the eighth started trading wildly. The firm reportedly lost about $440 million in under an hour. A repeatable pipeline that verifies every server, with a quick rollback, is the cure for “someone forgot one server”.

A minimal CodeBuild file that tests, builds and pushes an image:

# buildspec.yml
version: 0.2
phases:
  install:
    runtime-versions:
      nodejs: 20
  pre_build:
    commands:
      - npm ci
      - npm run lint
      - npm test                     # a failing test stops the pipeline here
      - aws ecr get-login-password --region $AWS_REGION | docker login --username AWS --password-stdin $ECR_URI
  build:
    commands:
      # tag with the commit id, never "latest", so you know exactly what runs
      - docker build -t $ECR_URI:$CODEBUILD_RESOLVED_SOURCE_VERSION .
      - docker push $ECR_URI:$CODEBUILD_RESOLVED_SOURCE_VERSION

Two habits matter more than the tool: tag images with the commit, and have a rollback plan (redeploy the previous tag). More on safe rollouts in section 5.


2. Infrastructure as Code: no click-ops

If your servers, buckets and databases were created by clicking in the console, nobody can review them, repeat them, or recreate them after a disaster. Infrastructure as Code (IaC) puts them in files, in Git, reviewed like application code. On AWS the native tool is CloudFormation (with CDK to write it in a real language); Terraform is the popular alternative.

Failure scenario (very common). A developer fixes a production problem by editing a security group in the console at 2 AM. It works. Weeks later the stack is redeployed from code, the manual fix disappears, and the same outage returns. The difference between what the code says and what is really running is called drift.

# template.yaml: a safe S3 bucket, defined once, reviewed in a pull request
Resources:
  UploadsBucket:
    Type: AWS::S3::Bucket
    Properties:
      VersioningConfiguration:
        Status: Enabled              # recover from accidental deletes
      BucketEncryption:
        ServerSideEncryptionConfiguration:
          - ServerSideEncryptionByDefault:
              SSEAlgorithm: AES256
      PublicAccessBlockConfiguration: # one setting prevents the classic "public bucket" leak
        BlockPublicAcls: true
        BlockPublicPolicy: true
        IgnorePublicAcls: true
        RestrictPublicBuckets: true

Rule of thumb: if it is not in code, it does not exist. Use CloudFormation’s drift detection to catch manual changes.


3. Secrets and config: keep them out of Git

Code is the same in every environment. Config (database host, feature settings) and secrets (passwords, API keys) change per environment and must live outside the code. On AWS use SSM Parameter Store for plain config and Secrets Manager for secrets (with automatic rotation).

Failure scenario. A developer commits an AWS access key to a public GitHub repository “just for testing”. Automated scanners look for exactly this and can find keys quickly, and attackers use them to launch expensive compute for crypto-mining. The bill arrives at the end of the month. (AWS also scans public repos and may quarantine exposed keys, but do not rely on that.)

// Read the secret at runtime. Nothing sensitive is stored in the repo or the image.
import {
  SecretsManagerClient,
  GetSecretValueCommand,
} from "@aws-sdk/client-secrets-manager";

const client = new SecretsManagerClient({ region: "us-east-1" });

export async function getDbPassword() {
  const res = await client.send(
    new GetSecretValueCommand({ SecretId: "prod/app/db" })
  );
  return JSON.parse(res.SecretString).password;
}

Notice there are no AWS keys in this code. When it runs on ECS or Lambda, the SDK automatically uses the IAM role of the service. That brings us to the most important topic.


4. IAM: give the least power possible

IAM (Identity and Access Management) decides who can do what. The principle of least privilege means each user or service gets only the permissions it needs, nothing more.

Failure scenario. In the 2019 Capital One breach, an attacker exploited a misconfigured web application firewall to make the server request its own instance metadata, which returned temporary credentials for the server’s IAM role. That role could read far more S3 data than the app ever needed, and data on about 100 million people in the US was copied. The first mistake opened the door; the over-broad role decided how much was behind it.

// BAD: the app can do anything to everything
{
  "Effect": "Allow",
  "Action": "s3:*",
  "Resource": "*"
}
// GOOD: only two actions, only one bucket
{
  "Effect": "Allow",
  "Action": ["s3:GetObject", "s3:PutObject"],
  "Resource": "arn:aws:s3:::my-app-uploads/*"
}

Quick IAM habits:

  • Use roles, not long-lived access keys, for services and CI.
  • Turn on MFA for humans and never use the root account for daily work.
  • Require IMDSv2 on EC2 (it blocks the simple metadata attack above).
  • Start narrow, and let IAM Access Analyzer show you what is unused.

5. Scaling and availability: assume things fail

Servers crash, disks fill up, a whole data centre can lose power. AWS Regions are split into Availability Zones (AZs), which are separate data centres. Spread your app across at least two, so losing one is boring, not an incident.

Users Application Load Balancer health checks hide unhealthy targets AWS Region Availability Zone A Availability Zone B App containers (ECS Fargate) Auto Scaling adds or removes tasks App containers (ECS Fargate) Auto Scaling adds or removes tasks RDS database: primary reads and writes RDS database: standby takes over automatically sync
A basic highly available setup. Lose one zone and the other keeps serving.

The building blocks:

  • Load balancer (ALB) spreads traffic and stops sending it to unhealthy servers.
  • Auto Scaling adds capacity when traffic rises and replaces broken instances.
  • RDS Multi-AZ keeps a standby database in another zone and fails over automatically.

Your app must cooperate. Give the load balancer an honest health endpoint:

// Liveness: "is the process up?" Readiness: "can it actually serve traffic?"
app.get("/health", async (req, res) => {
  try {
    await db.query("SELECT 1");      // can we reach the database?
    res.status(200).json({ status: "ok" });
  } catch (err) {
    res.status(503).json({ status: "db-unreachable" }); // ALB stops routing here
  }
});

And make rollouts safe. This ECS setting stops a bad release and rolls back on its own:

# CloudFormation: AWS::ECS::Service
DeploymentConfiguration:
  MinimumHealthyPercent: 100        # keep full capacity during the rollout
  MaximumPercent: 200               # start new tasks before stopping old ones
  DeploymentCircuitBreaker:
    Enable: true
    Rollback: true                  # new tasks keep failing? go back automatically

Failure scenario. Back to the 2017 S3 outage: many applications worked in one region only, so when that region’s S3 had trouble they went down with it. Multi-AZ protects you from a data centre failing. For truly critical systems, think about multi-region too, and about what your app does when a dependency is down (timeouts, retries, graceful errors).


6. Monitoring and alerting: know before your users do

You cannot fix what you cannot see. On AWS, CloudWatch gives you three things: logs (what happened), metrics (numbers over time such as CPU, 5xx errors, latency) and alarms (tell a human when a number looks wrong). X-Ray adds traces to follow one request across services.

Your app JSON logs, metrics CloudWatch stores + graphs Alarm 5xx > 10 for 3 min SNS email, Slack, pager You, on call read runbook, fix
From a failing request to a human who can act, without anyone staring at a dashboard.

Write logs that a machine can search. One JSON object per line beats a sentence:

// Searchable in CloudWatch Logs Insights: filter by orderId, sort by durationMs
console.log(JSON.stringify({
  level: "error",
  message: "payment failed",
  orderId: order.id,
  durationMs: Date.now() - start,
  requestId: req.headers["x-amzn-trace-id"], // follow one request end to end
}));

Then create an alarm so you hear about it first:

aws cloudwatch put-metric-alarm \
  --alarm-name api-5xx-high \
  --namespace AWS/ApplicationELB \
  --metric-name HTTPCode_Target_5XX_Count \
  --dimensions Name=LoadBalancer,Value=app/my-alb/50dc6c495c0c9188 \
  --statistic Sum --period 60 --evaluation-periods 3 \
  --threshold 10 --comparison-operator GreaterThanThreshold \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:oncall

Failure scenario (very common). A log file or database disk slowly fills up. Nobody has an alarm on disk space. On a Friday evening the service stops writing and falls over, and the first report is an angry customer email. A single alarm on “disk above 80%” turns that into a calm Tuesday task.

Alert on symptoms users feel (errors, latency) rather than every CPU blip, or people will start ignoring the pager.


7. Backups and recovery: test the restore

Two numbers decide your recovery plan. RPO (recovery point objective): how much data can you afford to lose? RTO (recovery time objective): how long can you afford to be down?

AWS gives you the tools: RDS automated backups with point-in-time restore, S3 versioning (shown in the template above), and AWS Backup to manage it all in one place.

Failure scenario. A bad migration deletes a column of customer data. The team is sure it has backups. Nobody ever tried restoring one, and the restore takes far longer than anyone expected. A backup you have never restored is only a hope.

Schedule a restore drill every few months: restore into a temporary database, check the data, and time it.


Quick map: concept to AWS service

Concept Why it matters AWS service
CI/CD Repeatable, safe releases CodePipeline, CodeBuild, CodeDeploy
Infrastructure as code Reviewable, rebuildable setup CloudFormation, CDK
Secrets and config Nothing sensitive in Git Secrets Manager, SSM Parameter Store
Access control Limit the damage of any mistake IAM, Access Analyzer
Scaling and availability Survive failures and traffic spikes ALB, Auto Scaling, Multi-AZ RDS
Monitoring Know before users do CloudWatch, X-Ray, SNS
Backup and recovery Survive data loss AWS Backup, RDS snapshots, S3 versioning

Your first-week checklist

  • Every deploy goes through a pipeline; images are tagged with the commit.
  • Infrastructure lives in Git (CloudFormation, CDK or Terraform).
  • No secrets in code, .env files in Git, or container images.
  • Each service has its own IAM role with only the permissions it needs.
  • MFA is on for every human, and the root account is locked away.
  • The app runs in at least two Availability Zones with a real health check.
  • Logs are structured, and an alarm exists for 5xx errors and disk space.
  • Backups are on, and you have restored one at least once.
  • A budget alert is set (AWS Budgets), so a forgotten resource cannot surprise you.

Key takeaways

  • DevOps is a habit, not a tool. Build, ship and run in small, automated, repeatable steps.
  • Automate deploys and keep a rollback ready; manual steps are where outages like Knight Capital’s start.
  • Put infrastructure in code so it can be reviewed and rebuilt, and watch for drift.
  • Keep secrets out of Git, and give every service least-privilege IAM so one mistake stays small.
  • Assume failure: use multiple Availability Zones, health checks and automatic rollbacks.
  • Monitor symptoms, alert a human, and test your restores.

You do not need to learn all of AWS. Pick one project this week, run through the checklist above, and fix the first gap you find. One safe step at a time is how real DevOps starts.