On 28 February 2017, an engineer at Amazon ran a routine command to remove a few servers from a billing system. One input was typed wrong, far more servers were removed than intended, and a large part of the internet that relied on S3 in the us-east-1 region went down for hours. AWS later described the cause in its public post-event summary and added safeguards so the same command could not remove that much capacity so quickly.
The lesson is not “engineers make typos”. The lesson is that systems must be built so a single mistake cannot become a disaster. That is what DevOps is about, and in 2026 every software engineer is expected to understand the basics, not only the “ops person”.
In this note you will learn:
- what DevOps really means, in one picture,
- seven basics (CI/CD, infrastructure as code, secrets, IAM, scaling, monitoring, backups),
- the AWS service behind each one,
- real failure scenarios and the code or config that prevents them.
Reading time: about 11 minutes.
What DevOps really means
DevOps is not a job title or a tool. It is a way of working where the people who build software also take part in running it, supported by automation. The goal: ship small changes often, safely, and know quickly when something breaks.
One more idea before we start: the shared responsibility model. AWS secures the data centres, hardware and core services. You are responsible for what you build on top: your code, your permissions, your configuration and your data. Almost every story below is a failure on the “you” side.
1. CI/CD: never deploy by hand
Continuous Integration (CI) means every change is automatically built and tested. Continuous Delivery (CD) means passing builds are deployed through the same repeatable steps every time. On AWS: CodePipeline orchestrates, CodeBuild builds and tests, CodeDeploy or ECS rolls out. (GitHub Actions works just as well and can deploy to AWS using an IAM role.)
Failure scenario. In 2012 the trading firm Knight Capital deployed new software by copying it to its servers by hand. By most public accounts it reached seven of eight servers, and old, dormant code on the eighth started trading wildly. The firm reportedly lost about $440 million in under an hour. A repeatable pipeline that verifies every server, with a quick rollback, is the cure for “someone forgot one server”.
A minimal CodeBuild file that tests, builds and pushes an image:
# buildspec.yml
version: 0.2
phases:
install:
runtime-versions:
nodejs: 20
pre_build:
commands:
- npm ci
- npm run lint
- npm test # a failing test stops the pipeline here
- aws ecr get-login-password --region $AWS_REGION | docker login --username AWS --password-stdin $ECR_URI
build:
commands:
# tag with the commit id, never "latest", so you know exactly what runs
- docker build -t $ECR_URI:$CODEBUILD_RESOLVED_SOURCE_VERSION .
- docker push $ECR_URI:$CODEBUILD_RESOLVED_SOURCE_VERSION
Two habits matter more than the tool: tag images with the commit, and have a rollback plan (redeploy the previous tag). More on safe rollouts in section 5.
2. Infrastructure as Code: no click-ops
If your servers, buckets and databases were created by clicking in the console, nobody can review them, repeat them, or recreate them after a disaster. Infrastructure as Code (IaC) puts them in files, in Git, reviewed like application code. On AWS the native tool is CloudFormation (with CDK to write it in a real language); Terraform is the popular alternative.
Failure scenario (very common). A developer fixes a production problem by editing a security group in the console at 2 AM. It works. Weeks later the stack is redeployed from code, the manual fix disappears, and the same outage returns. The difference between what the code says and what is really running is called drift.
# template.yaml: a safe S3 bucket, defined once, reviewed in a pull request
Resources:
UploadsBucket:
Type: AWS::S3::Bucket
Properties:
VersioningConfiguration:
Status: Enabled # recover from accidental deletes
BucketEncryption:
ServerSideEncryptionConfiguration:
- ServerSideEncryptionByDefault:
SSEAlgorithm: AES256
PublicAccessBlockConfiguration: # one setting prevents the classic "public bucket" leak
BlockPublicAcls: true
BlockPublicPolicy: true
IgnorePublicAcls: true
RestrictPublicBuckets: true
Rule of thumb: if it is not in code, it does not exist. Use CloudFormation’s drift detection to catch manual changes.
3. Secrets and config: keep them out of Git
Code is the same in every environment. Config (database host, feature settings) and secrets (passwords, API keys) change per environment and must live outside the code. On AWS use SSM Parameter Store for plain config and Secrets Manager for secrets (with automatic rotation).
Failure scenario. A developer commits an AWS access key to a public GitHub repository “just for testing”. Automated scanners look for exactly this and can find keys quickly, and attackers use them to launch expensive compute for crypto-mining. The bill arrives at the end of the month. (AWS also scans public repos and may quarantine exposed keys, but do not rely on that.)
// Read the secret at runtime. Nothing sensitive is stored in the repo or the image.
import {
SecretsManagerClient,
GetSecretValueCommand,
} from "@aws-sdk/client-secrets-manager";
const client = new SecretsManagerClient({ region: "us-east-1" });
export async function getDbPassword() {
const res = await client.send(
new GetSecretValueCommand({ SecretId: "prod/app/db" })
);
return JSON.parse(res.SecretString).password;
}
Notice there are no AWS keys in this code. When it runs on ECS or Lambda, the SDK automatically uses the IAM role of the service. That brings us to the most important topic.
4. IAM: give the least power possible
IAM (Identity and Access Management) decides who can do what. The principle of least privilege means each user or service gets only the permissions it needs, nothing more.
Failure scenario. In the 2019 Capital One breach, an attacker exploited a misconfigured web application firewall to make the server request its own instance metadata, which returned temporary credentials for the server’s IAM role. That role could read far more S3 data than the app ever needed, and data on about 100 million people in the US was copied. The first mistake opened the door; the over-broad role decided how much was behind it.
// BAD: the app can do anything to everything
{
"Effect": "Allow",
"Action": "s3:*",
"Resource": "*"
}
// GOOD: only two actions, only one bucket
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::my-app-uploads/*"
}
Quick IAM habits:
- Use roles, not long-lived access keys, for services and CI.
- Turn on MFA for humans and never use the root account for daily work.
- Require IMDSv2 on EC2 (it blocks the simple metadata attack above).
- Start narrow, and let IAM Access Analyzer show you what is unused.
5. Scaling and availability: assume things fail
Servers crash, disks fill up, a whole data centre can lose power. AWS Regions are split into Availability Zones (AZs), which are separate data centres. Spread your app across at least two, so losing one is boring, not an incident.
The building blocks:
- Load balancer (ALB) spreads traffic and stops sending it to unhealthy servers.
- Auto Scaling adds capacity when traffic rises and replaces broken instances.
- RDS Multi-AZ keeps a standby database in another zone and fails over automatically.
Your app must cooperate. Give the load balancer an honest health endpoint:
// Liveness: "is the process up?" Readiness: "can it actually serve traffic?"
app.get("/health", async (req, res) => {
try {
await db.query("SELECT 1"); // can we reach the database?
res.status(200).json({ status: "ok" });
} catch (err) {
res.status(503).json({ status: "db-unreachable" }); // ALB stops routing here
}
});
And make rollouts safe. This ECS setting stops a bad release and rolls back on its own:
# CloudFormation: AWS::ECS::Service
DeploymentConfiguration:
MinimumHealthyPercent: 100 # keep full capacity during the rollout
MaximumPercent: 200 # start new tasks before stopping old ones
DeploymentCircuitBreaker:
Enable: true
Rollback: true # new tasks keep failing? go back automatically
Failure scenario. Back to the 2017 S3 outage: many applications worked in one region only, so when that region’s S3 had trouble they went down with it. Multi-AZ protects you from a data centre failing. For truly critical systems, think about multi-region too, and about what your app does when a dependency is down (timeouts, retries, graceful errors).
6. Monitoring and alerting: know before your users do
You cannot fix what you cannot see. On AWS, CloudWatch gives you three things: logs (what happened), metrics (numbers over time such as CPU, 5xx errors, latency) and alarms (tell a human when a number looks wrong). X-Ray adds traces to follow one request across services.
Write logs that a machine can search. One JSON object per line beats a sentence:
// Searchable in CloudWatch Logs Insights: filter by orderId, sort by durationMs
console.log(JSON.stringify({
level: "error",
message: "payment failed",
orderId: order.id,
durationMs: Date.now() - start,
requestId: req.headers["x-amzn-trace-id"], // follow one request end to end
}));
Then create an alarm so you hear about it first:
aws cloudwatch put-metric-alarm \
--alarm-name api-5xx-high \
--namespace AWS/ApplicationELB \
--metric-name HTTPCode_Target_5XX_Count \
--dimensions Name=LoadBalancer,Value=app/my-alb/50dc6c495c0c9188 \
--statistic Sum --period 60 --evaluation-periods 3 \
--threshold 10 --comparison-operator GreaterThanThreshold \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:123456789012:oncall
Failure scenario (very common). A log file or database disk slowly fills up. Nobody has an alarm on disk space. On a Friday evening the service stops writing and falls over, and the first report is an angry customer email. A single alarm on “disk above 80%” turns that into a calm Tuesday task.
Alert on symptoms users feel (errors, latency) rather than every CPU blip, or people will start ignoring the pager.
7. Backups and recovery: test the restore
Two numbers decide your recovery plan. RPO (recovery point objective): how much data can you afford to lose? RTO (recovery time objective): how long can you afford to be down?
AWS gives you the tools: RDS automated backups with point-in-time restore, S3 versioning (shown in the template above), and AWS Backup to manage it all in one place.
Failure scenario. A bad migration deletes a column of customer data. The team is sure it has backups. Nobody ever tried restoring one, and the restore takes far longer than anyone expected. A backup you have never restored is only a hope.
Schedule a restore drill every few months: restore into a temporary database, check the data, and time it.
Quick map: concept to AWS service
| Concept | Why it matters | AWS service |
|---|---|---|
| CI/CD | Repeatable, safe releases | CodePipeline, CodeBuild, CodeDeploy |
| Infrastructure as code | Reviewable, rebuildable setup | CloudFormation, CDK |
| Secrets and config | Nothing sensitive in Git | Secrets Manager, SSM Parameter Store |
| Access control | Limit the damage of any mistake | IAM, Access Analyzer |
| Scaling and availability | Survive failures and traffic spikes | ALB, Auto Scaling, Multi-AZ RDS |
| Monitoring | Know before users do | CloudWatch, X-Ray, SNS |
| Backup and recovery | Survive data loss | AWS Backup, RDS snapshots, S3 versioning |
Your first-week checklist
- Every deploy goes through a pipeline; images are tagged with the commit.
- Infrastructure lives in Git (CloudFormation, CDK or Terraform).
- No secrets in code,
.envfiles in Git, or container images. - Each service has its own IAM role with only the permissions it needs.
- MFA is on for every human, and the root account is locked away.
- The app runs in at least two Availability Zones with a real health check.
- Logs are structured, and an alarm exists for 5xx errors and disk space.
- Backups are on, and you have restored one at least once.
- A budget alert is set (AWS Budgets), so a forgotten resource cannot surprise you.
Key takeaways
- DevOps is a habit, not a tool. Build, ship and run in small, automated, repeatable steps.
- Automate deploys and keep a rollback ready; manual steps are where outages like Knight Capital’s start.
- Put infrastructure in code so it can be reviewed and rebuilt, and watch for drift.
- Keep secrets out of Git, and give every service least-privilege IAM so one mistake stays small.
- Assume failure: use multiple Availability Zones, health checks and automatic rollbacks.
- Monitor symptoms, alert a human, and test your restores.
You do not need to learn all of AWS. Pick one project this week, run through the checklist above, and fix the first gap you find. One safe step at a time is how real DevOps starts.