Intentional Release Guidelines
This document establishes the principles and policies governing how we build, version, and deploy releases. It is intended for all engineers and team leads involved in the release process.
For the step-by-step workflow walkthrough (branching, CI pipelines, Kargo promotion), see the Intentional Release Workflow Guide. For detailed database migration patterns (Expand-Contract implementations, code examples), see the Database Migration Patterns Reference.
Why We’re Changing
Our current deployment process deploys whatever has accumulated on main, either continuously or via daily scheduled pushes, without an intentional release decision. This creates a set of compounding problems:
We deploy blind. There is low visibility into what’s being deployed until it’s already in production. Correlating a production issue with the specific change that caused it requires manually diffing commit SHAs. There is no changelog, no release documentation, and no structured record of what went out.
Database migrations go unassessed. Daily production deployments may include destructive schema changes without warning. We have no pre-deployment analysis of migration risk, and in most cases we don’t have a process, or even the ability, to roll back migrations after they run.
Staging doesn’t validate functionality. Our staging environment verifies that a deployment succeeds mechanically, but does not test whether the application actually works. Production is frequently the first environment where we discover functional issues.
Rollbacks are unreliable or impossible. When something goes wrong, we can’t safely roll back because we don’t immediately know what migrations ran, whether they’re reversible, and the previous code version doesn’t contain the down migrations for the newer changes. We either intervene manually or roll forward.
These problems are addressed by shifting from continuous deployment to intentional, tag-based releases with structured documentation, automated risk analysis, and environment-based promotion.
Core Principles
1. Deploy from Tags, Not Commits
We stop deploying directly from main branch commits. Instead, we create tagged releases (v1.2.3) that represent intentional deployment decisions. Each release contains a curated and documented set of changes, not an accidental combination of whatever merged since the last deploy.
The main branch streams continuously to the Dev environment for smoke testing. Deploying to Staging and Production requires a tagged release.
2. Version Numbers Communicate Risk
We use a versioning scheme inspired by semantic versioning, but adapted for application deployment rather than library publishing. In standard semver, version numbers communicate API compatibility to external consumers. We are not publishing libraries, so we repurpose the MAJOR/MINOR/PATCH levels to communicate deployment risk to the team performing the release, specifically whether the release includes database migrations and what kind.
PATCH (v1.2.3 → v1.2.4): Minimal risk. Bug fixes, configuration updates. No database changes whatsoever. Safe to deploy and roll back at any time.
MINOR (v1.2.4 → v1.3.0): Low risk. New features with additive-only database changes (new tables, new columns). New API endpoints that don’t break existing functionality. Rollback is possible, though new feature data may be lost.
MAJOR (v1.5.2 → v2.0.0): High risk. Breaking API changes, or database migrations that modify or remove existing structures. Rollback requires planning, may involve data loss, or may be impossible. These releases require an Expand-Contract migration strategy.
3. Every Release Is Documented
Every release must include structured documentation that describes what changed, what the deployment risk is, what the rollback strategy is, and what service dependencies exist. This documentation is generated automatically where possible (via git-cliff for changelogs, CI analysis for migration risk) and supplemented manually by the releasing team.
4. Database Rollbacks Are Not a Safety Net
Down migrations rarely work in production. Failed migrations leave databases in unknown states. Data written after deployment cannot be “un-written.” We plan forward-only migrations and use the Expand-Contract pattern for breaking schema changes. Version numbers and release documentation make the risk visible so teams can make informed decisions before deploying, not after.
5. Releases Are Promoted Through Environments, Not Pushed to Production
Releases follow a promotion path: Dev → Staging → Production. Each environment serves a distinct purpose with different entry criteria. Promotion between environments is controlled by Kargo and requires validation to pass before a release advances.
Release Workflow Overview
The release lifecycle follows this path:
flowchart TD
subgraph DEV["DEVELOPMENT"]
A[Feature Branch] -->|Merge Request| B[Merge to main]
end
subgraph CI["CI ON MERGE TO MAIN"]
C[Build & test] --> D["Push Docker image<br/><i>tagged with commit SHA</i>"]
D --> E["Stream to Dev environment<br/><i>continuous</i>"]
D --> F[Run pre-release analysis]
F --> F1["Generate changelog<br/><i>git-cliff</i>"]
F --> F2["Detect migrations<br/>& classify risk"]
F --> F3["Recommend version bump<br/><i>PATCH / MINOR / MAJOR</i>"]
F1 & F2 & F3 --> G["Create / update<br/>draft release"]
end
subgraph DECISION["RELEASE DECISION · Manual"]
H["Team reviews draft release:<br/>changelog, migration risk,<br/>version bump, dependencies"]
H -->|Publish release| I(["Tag: v1.3.0"])
end
subgraph PROMOTION["ENVIRONMENT PROMOTION"]
J["Staging<br/><i>Kargo detects new tag</i>"]
J --> K{"Validation passes?<br/><i>Analysis template:<br/>migration safety,<br/>smoke tests</i>"}
K -->|Yes| L["Production<br/><i>Manual approval required</i>"]
K -->|No| M["Release blocked<br/>Team investigates"]
end
B --> C
G --> H
I --> J
style A fill:#e8eef7,stroke:#4a6fa5,color:#000
style B fill:#e8eef7,stroke:#4a6fa5,color:#000
style C fill:#e8eef7,stroke:#4a6fa5,color:#000
style D fill:#e8eef7,stroke:#4a6fa5,color:#000
style E fill:#e8eef7,stroke:#4a6fa5,color:#000
style F fill:#e8eef7,stroke:#4a6fa5,color:#000
style F1 fill:#e8eef7,stroke:#4a6fa5,color:#000
style F2 fill:#e8eef7,stroke:#4a6fa5,color:#000
style F3 fill:#e8eef7,stroke:#4a6fa5,color:#000
style G fill:#e8eef7,stroke:#4a6fa5,color:#000
style H fill:#fdf3e3,stroke:#d4a043,color:#000
style I fill:#ffc107,stroke:#d4a043,color:#000
style J fill:#e6f4ea,stroke:#28a745,color:#000
style K fill:#e6f4ea,stroke:#28a745,color:#000
style L fill:#e6f4ea,stroke:#28a745,color:#000
style M fill:#f8d7da,stroke:#dc3545,color:#000
Key points:
- CI runs pre-release analysis on every merge to
main, but does not create a release. It produces a draft. - The team makes an explicit decision to publish a release and tag it. This is the “intentional” part.
- Staging only receives tagged releases, not raw main-branch commits.
- Production only receives releases that have been validated in Staging.
Version Bump Decision Guide
Use this decision tree to determine the correct version bump for a release. The CI pipeline will recommend a bump automatically, but the releasing team is responsible for the final decision.
flowchart TD
A["What changes are in this release?"] --> B{"Any database<br/>migrations?"}
B -- No --> C{"Any breaking<br/>API changes?"}
C -- No --> PATCH["<b>PATCH</b><br/>Bug fixes, config updates<br/>Safe to roll back anytime"]
C -- Yes --> MAJOR1["<b>MAJOR</b><br/>Breaking API contract changes<br/>Coordinate with consumers"]
B -- Yes --> D{"What kind<br/>of migration?"}
D -- "Additive only<br/>(CREATE TABLE, ADD COLUMN)" --> E{"Any breaking<br/>API changes?"}
E -- No --> MINOR["<b>MINOR</b><br/>New features, additive schema<br/>Rollback loses new data only"]
E -- Yes --> MAJOR2["<b>MAJOR</b><br/>Breaking API + additive schema<br/>Coordinate with consumers"]
D -- "Destructive or modifying<br/>(DROP, ALTER type, remove constraint)" --> MAJOR3["<b>MAJOR</b><br/>Requires Expand-Contract strategy<br/>Rollback requires planning"]
style A fill:#e8eef7,stroke:#4a6fa5,color:#000
style B fill:#e8eef7,stroke:#4a6fa5,color:#000
style C fill:#e8eef7,stroke:#4a6fa5,color:#000
style D fill:#e8eef7,stroke:#4a6fa5,color:#000
style E fill:#e8eef7,stroke:#4a6fa5,color:#000
style PATCH fill:#d4edda,stroke:#28a745,color:#000
style MINOR fill:#fff3cd,stroke:#d4a043,color:#000
style MAJOR1 fill:#f8d7da,stroke:#dc3545,color:#000
style MAJOR2 fill:#f8d7da,stroke:#dc3545,color:#000
style MAJOR3 fill:#f8d7da,stroke:#dc3545,color:#000
Rules of thumb:
- If in doubt between MINOR and MAJOR, choose MAJOR. It’s better to over-communicate risk.
- Renaming a column is MAJOR. It’s a DROP + ADD from the database’s perspective, and existing code will break.
- Additive migrations (
CREATE TABLE,ADD COLUMNwith nullable or safe defaults,CREATE INDEX CONCURRENTLY) are MINOR. Old code continues to work because it simply doesn’t reference the new structures. - Constraint additions on populated columns are MAJOR. They can lock tables and fail on existing data.
- Avoid combining breaking API changes and destructive migrations in the same MAJOR release when possible. Each carries distinct risk and requires different coordination. Keeping them separate makes rollback reasoning and cross-team communication simpler.
Environment Strategy
Our deployment infrastructure has three environments, each with a distinct purpose:
Note: During the transition to GitOps, deployments are triggered via CI/CD pipelines. The table below describes the target state.
| Environment | Source | Deployment Trigger | Purpose |
|---|---|---|---|
| Dev | main branch (HEAD) |
Automatic on every merge | Deployment verification. Engineers confirm that changes build, deploy, and run correctly. Available for light manual checks, but less controlled than Staging since it receives every merge continuously. |
| Staging | Tagged releases only | Kargo detects new tag | Release validation. Kargo Analysis templates run migration safety checks and smoke tests. Releases must pass validation here before they are eligible for Production. |
| Production | Validated Staging releases | Manual approval (recommended) or automatic after Staging validation | Live environment. Only receives releases that passed Staging validation. |
What this changes from today:
- Today, staging receives whatever is on
mainand only tests deployment mechanics. In the new model, Staging receives intentional tagged releases and validates them functionally. - The new Dev environment takes over what Staging does today: continuously deploying from
mainto verify that applications build and run correctly. This frees Staging to become a real QA environment focused on validating tagged releases before promotion. - Today, production deployments happen automatically or on a daily schedule. In the new model, promotion to Production requires reviewing and approving a release that has been validated in Staging. Automatic promotion after successful Staging validation is possible, but manual approval is recommended so teams can review the release notes and migration risk before going live.
Production approval ownership: Manual approval for Production promotion is the responsibility of the releasing team (staff engineer, team lead or manager). There is no strict role requirement at this stage; the key expectation is that someone on the team has reviewed the release notes, migration risk, and rollback strategy before approving.
Release Documentation Standards
Every published release must include the following metadata. The CI pipeline generates most of this automatically via git-cliff (changelog) and migration analysis tooling. The releasing team adds or corrects anything the automation misses.
Release: v2.1.0
Type: MINOR
Changelog:
Features:
- "[#123] Add user email preferences"
- "[#456] Support bulk export from dashboard"
Fixes:
- "[#789] Fix password validation edge case on special characters"
Database Changes:
Migrations:
- "Added user_preferences table"
- "Added email_verified column to users table (nullable, no default)"
Risk Level: LOW (additive only)
Rollback Impact: "New user_preferences data would be lost if rolled back"
API Changes:
New Endpoints:
- "POST /api/user/preferences"
Breaking Changes: None
Service Dependencies:
Requires: "frontend >= v2.1.0 (uses new preferences API)"
Compatible With: "All existing API consumers (no breaking changes)"
Rollback Strategy:
Code: "Safe to roll back to v2.0.x"
Database: "New tables/columns can be left in place; no destructive rollback needed"
Data Impact: "User preferences created after deployment would be orphaned"What gets automated vs. what’s manual:
| Field | Source |
|---|---|
| Changelog / Commits | Automated via git-cliff from conventional commits |
| Database migration detection | Automated by CI scanning migration directories |
| Migration risk classification | Automated by analysis tooling (Squawk) |
| Version bump recommendation | Automated based on migration type + commit prefixes |
| API changes | Manual: team documents new endpoints and breaking changes |
| Service dependencies | Manual: team documents cross-service requirements |
| Rollback strategy narrative | Manual: team assesses based on migration risk and data impact |
Database Migration Policy
The Forward-Only Principle
Database rollbacks are unreliable in production for three fundamental reasons:
Partial failures leave unknown state. If an “up” migration adds two columns and fails after the first, the database is in a state the “down” migration doesn’t expect. Running the down migration will fail, leaving you stuck.
Reversing additive changes destroys data. Dropping a column that was successfully added doesn’t “undo” the addition. It permanently deletes all data in that column. Re-running the migration won’t restore it.
Previous code doesn’t contain future down migrations. When you roll back to a previous application version, that version’s codebase doesn’t include the down migrations for the changes that came after it. The rollback artifacts simply don’t exist in the image you’re deploying.
For these reasons, we treat database migrations as forward-only operations and use the Expand-Contract pattern for breaking schema changes.
Migration Risk Classifications
The CI pipeline classifies every detected migration and recommends a version bump based on the Version Bump Decision Guide above.
The Expand-Contract Pattern
Any schema change that would break the currently-running application code requires the Expand-Contract pattern. The core idea is to split one breaking change into a series of backward-compatible steps, each deployed as its own release. Because each step is backward-compatible, it can be rolled back independently without data loss (until the final cleanup step).
The pattern works in five steps:
Step 1: Expand. Add the new structure (columns, tables) alongside the old. Update application code to write to both old and new structures, but continue reading from the old. At this point the new structure exists but is not yet trusted. This step is a MINOR release — it’s purely additive. Rolling back means dropping the new structure, which is safe because nothing depends on it yet.
Step 2: Backfill. Migrate existing data from the old structure into the new. This runs as a background task (batched updates, throttled to avoid overwhelming the database). The application still reads from the old structure, so this step doesn’t affect correctness. Rolling back means clearing the new structure and restarting the backfill.
Step 3: Switch reads. Once the backfill is complete and data exists consistently in both structures, switch the application to read from the new structure while still writing to both. This is the key transition point. Rolling back means switching reads back to the old structure, which still has consistent data because dual-writes are still active.
Step 4: Stop writing to old structure. The old structure is now neither read from nor written to. This step is harder to roll back because the old structure stops receiving updates. However, there should be little need to revert since the new structure is already fully validated.
Step 5: Contract. Remove the old structure. This is the only truly irreversible step — it’s a MAJOR release. It should only happen after the new structure has been running in production long enough to be confident it’s correct.
How this maps to releases: Steps 1 through 4 are each a MINOR release (backward-compatible, additive or neutral). Step 5 is the only MAJOR release. If anything goes wrong at Steps 1-3, you can roll back that step’s release safely. The risk is concentrated in Step 5, which only runs after everything else is validated.
For implementation details including Rails and Python code examples, see the Database Migration Patterns Reference.
Migration Analysis in CI
Every merge to main triggers migration analysis that feeds into the draft release (pre-release):
- Rails and Python services: Migration DSLs are converted to raw SQL, which is then linted by Squawk in CI.
- Output: The analysis produces a risk classification, a recommended version bump, and any warnings (e.g., “this ALTER may lock the table for extended periods on tables over 1M rows”). This output is included in the draft release for the team to review.
The analysis is advisory, not blocking. It informs the team’s release decision rather than preventing merges or deploys.
Multi-Service Coordination
Terminology: Throughout this document, “service” refers to what we sometimes call an “application”. We use “service” here to emphasize the inter-service communication perspective.
Our architecture means each service maintains its own database and communicates with other services via APIs and message queues. Teams deploy independently on their own schedule. This works well because of our long-standing practice of maintaining backward compatibility across service boundaries.
Backward Compatibility Across Service Boundaries
We maintain backward compatibility on all inter-service contracts as our primary coordination mechanism. This applies to REST APIs, SNS/SQS message schemas, and any other interface between services. The key policies are:
New endpoints, new optional fields, and new message attributes are always safe. These are additive changes that don’t affect existing consumers. They map to MINOR releases.
Removing or renaming fields, changing response/message shapes, or adding required parameters are breaking changes. These require coordinated deployment across producer and consumer services, and they map to MAJOR releases on the producing service.
Deprecation before removal. When a field, endpoint, or message attribute needs to go away, mark it deprecated in one release, communicate the timeline to consuming teams, and remove it in a future release after consumers have migrated.
Coordinating Cross-Service Releases
When a release on one service requires a corresponding change on another, coordination relies on release documentation and team communication:
- The releasing team lists dependent services (and minimum compatible versions) under Service Dependencies in the release notes, making the requirement visible during review.
- Teams coordinate deployment ordering. Typically the backward-compatible side deploys first. For example: deploy the new API endpoint on Service A, then deploy the consumer code on Service B that uses it.
Configuration Coordination
Each service has its own repository containing application code and Dockerfile. Kubernetes deployment configuration (manifests, Kustomize overlays, environment-specific settings) lives separately in the deployment-config repository.
Because application code and deployment configuration are released independently, a complete deployment is the combination of a service image version and a deployment configuration version.
When a release requires a corresponding configuration change (e.g., new environment variables, updated resource limits), both should be documented in the release notes and coordinated during promotion.
Hotfix and Emergency Releases
When Production is broken and a fix is urgent, the standard release flow still applies, but with a compressed timeline. The goal is to keep the process intentional even under pressure, because emergencies are exactly when undocumented, unreviewed changes cause the most damage.
The key policies for emergency releases:
- The fix must land on
mainbefore the hotfix branch is closed. This guarantees the fix is part of mainline history and will not be lost in future releases. Ifmainis close to the broken tag, fixmainfirst and cherry-pick to a hotfix branch. Ifmainhas diverged significantly, fix the hotfix branch first, deploy, close the incident — then port tomainimmediately after. Do not let conflict resolution onmaindelay a production fix. - Hotfix releases should always be PATCH releases. If the fix itself requires a database migration or a breaking change, that is a sign the underlying issue needs a more considered approach rather than an emergency patch.
- Staging validation can be expedited but should not be skipped. Deploying an untagged commit directly to Production undermines traceability, which is exactly what you need most during an incident.
- Tagging, release documentation, and Staging promotion are never skipped. Even under time pressure, these are what make the fix traceable and auditable.
For the step-by-step hotfix workflow (branching, cherry-picking, expedited promotion), see Hotfix and Emergency Releases in the Workflow Guide.
What qualifies as a hotfix
Before compressing the release timeline, the first question is whether the situation actually warrants it. The standard release process exists for a reason — skipping steps under false urgency is how undocumented changes accumulate.
A hotfix is warranted when something is broken in production right now and cannot wait until the next business day without meaningful user or business impact. The bar is intentionally high.
flowchart TD
A([After-hours deploy needed?]) --> B{Is something broken\nin production right now?}
B -- No --> C[Not a hotfix.\nWait for business hours\nor the normal release process.]
B -- Yes --> D{Can it wait\nuntil tomorrow morning?}
D -- Yes --> E[It waits.\nSchedule for the next business day.]
D -- No --> F{Manager aware\nand signed off?}
F -- No --> G[Stop.\nContact your manager\nbefore proceeding.]
F -- Yes --> H{Incident opened +\nrollback plan documented?}
H -- No --> I[Open incident first.\nDocument rollback plan.\nThen deploy.]
H -- Yes --> J([✅ Proceed with compressed release flow])
style C fill:#f5f5f5,stroke:#ccc
style E fill:#f5f5f5,stroke:#ccc
style G fill:#fdecea,stroke:#e57373
style I fill:#fff8e1,stroke:#ffb300
style J fill:#e8f5e9,stroke:#66bb6a
Situations that do not qualify as hotfixes:
- A feature that was recently shipped and has a minor gap or missing behaviour
- Content availability issues affecting a small number of users on a low-traffic product
- Anything that can reasonably wait until the next business day
When in doubt, ask your manager. A five-minute conversation before a deploy is always worth it.
If you are deploying after hours, it is an incident
Any after-hours production deploy — regardless of perceived risk — must be opened as an incident in the incident management tool before the deploy happens. This is not bureaucracy. It is the mechanism that creates visibility, traceability, and a rollback record. Without it, there is no paper trail if something goes wrong, and no way for other teams to know what changed.
Minimum bar for any after-hours deploy:
| Requirement | Why |
|---|---|
| Incident opened in incident management tool | Creates a traceable record before any change is made |
| MR linked to a bug report or incident — not just the original feature | Ensures the fix is independently trackable |
| One-line Slack message in the relevant channel before deploying | No one should find out from the deployment itself |
| Manager aware and signed off | No surprises. Awareness is the floor; sign-off is expected for non-trivial risk |
| Rollback plan documented, even in one sentence | Forces the question before the deploy, not after |
Shared systems
For production systems with multiple teams deploying to them, an undocumented after-hours deploy by one team creates risk for all of them. The policies above apply regardless of which team owns the change. If your deploy touches a shared system, the manager awareness requirement extends to the leads of other affected teams.
Success Metrics
These metrics help us evaluate whether the transition to intentional releases is achieving its goals:
Release documentation coverage: Do releases consistently include the required metadata (changelog, migration risk, dependencies, rollback strategy)? Measured by CI validation on release creation.
Migration success rate: How often do database migrations succeed without requiring manual intervention? Measured by tracking migration failures in Staging and Production.
Deployment visibility: Can any engineer explain what changed in a given deployment by reading the release documentation? Validated through periodic team check-ins.
Rollback clarity: Before Production deployments, can the team articulate the rollback strategy and its limitations? This is part of the approval step.
Mean time to recovery (MTTR): Does the new process improve our ability to recover from deployment issues, either through faster rollbacks (for PATCH releases) or through better-informed roll-forward decisions (for MINOR/MAJOR releases)?