Skip to content
Intentional Release Guidelines

Intentional Release Guidelines

This document establishes the principles and policies governing how we build, version, and deploy releases. It is intended for all engineers and team leads involved in the release process.

For the step-by-step workflow walkthrough (branching, CI pipelines, Kargo promotion), see the Intentional Release Workflow Guide. For detailed database migration patterns (Expand-Contract implementations, code examples), see the Database Migration Patterns Reference.


Why We’re Changing

Our current deployment process deploys whatever has accumulated on main, either continuously or via daily scheduled pushes, without an intentional release decision. This creates a set of compounding problems:

We deploy blind. There is low visibility into what’s being deployed until it’s already in production. Correlating a production issue with the specific change that caused it requires manually diffing commit SHAs. There is no changelog, no release documentation, and no structured record of what went out.

Database migrations go unassessed. Daily production deployments may include destructive schema changes without warning. We have no pre-deployment analysis of migration risk, and in most cases we don’t have a process, or even the ability, to roll back migrations after they run.

Staging doesn’t validate functionality. Our staging environment verifies that a deployment succeeds mechanically, but does not test whether the application actually works. Production is frequently the first environment where we discover functional issues.

Rollbacks are unreliable or impossible. When something goes wrong, we can’t safely roll back because we don’t immediately know what migrations ran, whether they’re reversible, and the previous code version doesn’t contain the down migrations for the newer changes. We either intervene manually or roll forward.

These problems are addressed by shifting from continuous deployment to intentional, tag-based releases with structured documentation, automated risk analysis, and environment-based promotion.


Core Principles

1. Deploy from Tags, Not Commits

We stop deploying directly from main branch commits. Instead, we create tagged releases (v1.2.3) that represent intentional deployment decisions. Each release contains a curated and documented set of changes, not an accidental combination of whatever merged since the last deploy.

The main branch streams continuously to the Dev environment for smoke testing. Deploying to Staging and Production requires a tagged release.

2. Version Numbers Communicate Risk

We use a versioning scheme inspired by semantic versioning, but adapted for application deployment rather than library publishing. In standard semver, version numbers communicate API compatibility to external consumers. We are not publishing libraries, so we repurpose the MAJOR/MINOR/PATCH levels to communicate deployment risk to the team performing the release, specifically whether the release includes database migrations and what kind.

PATCH (v1.2.3 → v1.2.4): Minimal risk. Bug fixes, configuration updates. No database changes whatsoever. Safe to deploy and roll back at any time.

MINOR (v1.2.4 → v1.3.0): Low risk. New features with additive-only database changes (new tables, new columns). New API endpoints that don’t break existing functionality. Rollback is possible, though new feature data may be lost.

MAJOR (v1.5.2 → v2.0.0): High risk. Breaking API changes, or database migrations that modify or remove existing structures. Rollback requires planning, may involve data loss, or may be impossible. These releases require an Expand-Contract migration strategy.

3. Every Release Is Documented

Every release must include structured documentation that describes what changed, what the deployment risk is, what the rollback strategy is, and what service dependencies exist. This documentation is generated automatically where possible (via git-cliff for changelogs, CI analysis for migration risk) and supplemented manually by the releasing team.

4. Database Rollbacks Are Not a Safety Net

Down migrations rarely work in production. Failed migrations leave databases in unknown states. Data written after deployment cannot be “un-written.” We plan forward-only migrations and use the Expand-Contract pattern for breaking schema changes. Version numbers and release documentation make the risk visible so teams can make informed decisions before deploying, not after.

5. Releases Are Promoted Through Environments, Not Pushed to Production

Releases follow a promotion path: Dev → Staging → Production. Each environment serves a distinct purpose with different entry criteria. Promotion between environments is controlled by Kargo and requires validation to pass before a release advances.


Release Workflow Overview

The release lifecycle follows this path:

    flowchart TD
    subgraph DEV["DEVELOPMENT"]
        A[Feature Branch] -->|Merge Request| B[Merge to main]
    end

    subgraph CI["CI ON MERGE TO MAIN"]
        C[Build & test] --> D["Push Docker image<br/><i>tagged with commit SHA</i>"]
        D --> E["Stream to Dev environment<br/><i>continuous</i>"]
        D --> F[Run pre-release analysis]
        F --> F1["Generate changelog<br/><i>git-cliff</i>"]
        F --> F2["Detect migrations<br/>& classify risk"]
        F --> F3["Recommend version bump<br/><i>PATCH / MINOR / MAJOR</i>"]
        F1 & F2 & F3 --> G["Create / update<br/>draft release"]
    end

    subgraph DECISION["RELEASE DECISION · Manual"]
        H["Team reviews draft release:<br/>changelog, migration risk,<br/>version bump, dependencies"]
        H -->|Publish release| I(["Tag: v1.3.0"])
    end

    subgraph PROMOTION["ENVIRONMENT PROMOTION"]
        J["Staging<br/><i>Kargo detects new tag</i>"]
        J --> K{"Validation passes?<br/><i>Analysis template:<br/>migration safety,<br/>smoke tests</i>"}
        K -->|Yes| L["Production<br/><i>Manual approval required</i>"]
        K -->|No| M["Release blocked<br/>Team investigates"]
    end

    B --> C
    G --> H
    I --> J

    style A fill:#e8eef7,stroke:#4a6fa5,color:#000
    style B fill:#e8eef7,stroke:#4a6fa5,color:#000
    style C fill:#e8eef7,stroke:#4a6fa5,color:#000
    style D fill:#e8eef7,stroke:#4a6fa5,color:#000
    style E fill:#e8eef7,stroke:#4a6fa5,color:#000
    style F fill:#e8eef7,stroke:#4a6fa5,color:#000
    style F1 fill:#e8eef7,stroke:#4a6fa5,color:#000
    style F2 fill:#e8eef7,stroke:#4a6fa5,color:#000
    style F3 fill:#e8eef7,stroke:#4a6fa5,color:#000
    style G fill:#e8eef7,stroke:#4a6fa5,color:#000
    style H fill:#fdf3e3,stroke:#d4a043,color:#000
    style I fill:#ffc107,stroke:#d4a043,color:#000
    style J fill:#e6f4ea,stroke:#28a745,color:#000
    style K fill:#e6f4ea,stroke:#28a745,color:#000
    style L fill:#e6f4ea,stroke:#28a745,color:#000
    style M fill:#f8d7da,stroke:#dc3545,color:#000
  

Key points:

  • CI runs pre-release analysis on every merge to main, but does not create a release. It produces a draft.
  • The team makes an explicit decision to publish a release and tag it. This is the “intentional” part.
  • Staging only receives tagged releases, not raw main-branch commits.
  • Production only receives releases that have been validated in Staging.

Version Bump Decision Guide

Use this decision tree to determine the correct version bump for a release. The CI pipeline will recommend a bump automatically, but the releasing team is responsible for the final decision.

    flowchart TD
    A["What changes are in this release?"] --> B{"Any database<br/>migrations?"}

    B -- No --> C{"Any breaking<br/>API changes?"}
    C -- No --> PATCH["<b>PATCH</b><br/>Bug fixes, config updates<br/>Safe to roll back anytime"]
    C -- Yes --> MAJOR1["<b>MAJOR</b><br/>Breaking API contract changes<br/>Coordinate with consumers"]

    B -- Yes --> D{"What kind<br/>of migration?"}

    D -- "Additive only<br/>(CREATE TABLE, ADD COLUMN)" --> E{"Any breaking<br/>API changes?"}
    E -- No --> MINOR["<b>MINOR</b><br/>New features, additive schema<br/>Rollback loses new data only"]
    E -- Yes --> MAJOR2["<b>MAJOR</b><br/>Breaking API + additive schema<br/>Coordinate with consumers"]

    D -- "Destructive or modifying<br/>(DROP, ALTER type, remove constraint)" --> MAJOR3["<b>MAJOR</b><br/>Requires Expand-Contract strategy<br/>Rollback requires planning"]

    style A fill:#e8eef7,stroke:#4a6fa5,color:#000
    style B fill:#e8eef7,stroke:#4a6fa5,color:#000
    style C fill:#e8eef7,stroke:#4a6fa5,color:#000
    style D fill:#e8eef7,stroke:#4a6fa5,color:#000
    style E fill:#e8eef7,stroke:#4a6fa5,color:#000
    style PATCH fill:#d4edda,stroke:#28a745,color:#000
    style MINOR fill:#fff3cd,stroke:#d4a043,color:#000
    style MAJOR1 fill:#f8d7da,stroke:#dc3545,color:#000
    style MAJOR2 fill:#f8d7da,stroke:#dc3545,color:#000
    style MAJOR3 fill:#f8d7da,stroke:#dc3545,color:#000
  

Rules of thumb:

  • If in doubt between MINOR and MAJOR, choose MAJOR. It’s better to over-communicate risk.
  • Renaming a column is MAJOR. It’s a DROP + ADD from the database’s perspective, and existing code will break.
  • Additive migrations (CREATE TABLE, ADD COLUMN with nullable or safe defaults, CREATE INDEX CONCURRENTLY) are MINOR. Old code continues to work because it simply doesn’t reference the new structures.
  • Constraint additions on populated columns are MAJOR. They can lock tables and fail on existing data.
  • Avoid combining breaking API changes and destructive migrations in the same MAJOR release when possible. Each carries distinct risk and requires different coordination. Keeping them separate makes rollback reasoning and cross-team communication simpler.

Environment Strategy

Our deployment infrastructure has three environments, each with a distinct purpose:

Note: During the transition to GitOps, deployments are triggered via CI/CD pipelines. The table below describes the target state.

Environment Source Deployment Trigger Purpose
Dev main branch (HEAD) Automatic on every merge Deployment verification. Engineers confirm that changes build, deploy, and run correctly. Available for light manual checks, but less controlled than Staging since it receives every merge continuously.
Staging Tagged releases only Kargo detects new tag Release validation. Kargo Analysis templates run migration safety checks and smoke tests. Releases must pass validation here before they are eligible for Production.
Production Validated Staging releases Manual approval (recommended) or automatic after Staging validation Live environment. Only receives releases that passed Staging validation.

What this changes from today:

  • Today, staging receives whatever is on main and only tests deployment mechanics. In the new model, Staging receives intentional tagged releases and validates them functionally.
  • The new Dev environment takes over what Staging does today: continuously deploying from main to verify that applications build and run correctly. This frees Staging to become a real QA environment focused on validating tagged releases before promotion.
  • Today, production deployments happen automatically or on a daily schedule. In the new model, promotion to Production requires reviewing and approving a release that has been validated in Staging. Automatic promotion after successful Staging validation is possible, but manual approval is recommended so teams can review the release notes and migration risk before going live.

Production approval ownership: Manual approval for Production promotion is the responsibility of the releasing team (staff engineer, team lead or manager). There is no strict role requirement at this stage; the key expectation is that someone on the team has reviewed the release notes, migration risk, and rollback strategy before approving.


Release Documentation Standards

Every published release must include the following metadata. The CI pipeline generates most of this automatically via git-cliff (changelog) and migration analysis tooling. The releasing team adds or corrects anything the automation misses.

Release: v2.1.0
Type: MINOR

Changelog:
  Features:
    - "[#123] Add user email preferences"
    - "[#456] Support bulk export from dashboard"
  Fixes:
    - "[#789] Fix password validation edge case on special characters"

Database Changes:
  Migrations:
    - "Added user_preferences table"
    - "Added email_verified column to users table (nullable, no default)"
  Risk Level: LOW (additive only)
  Rollback Impact: "New user_preferences data would be lost if rolled back"

API Changes:
  New Endpoints:
    - "POST /api/user/preferences"
  Breaking Changes: None

Service Dependencies:
  Requires: "frontend >= v2.1.0 (uses new preferences API)"
  Compatible With: "All existing API consumers (no breaking changes)"

Rollback Strategy:
  Code: "Safe to roll back to v2.0.x"
  Database: "New tables/columns can be left in place; no destructive rollback needed"
  Data Impact: "User preferences created after deployment would be orphaned"

What gets automated vs. what’s manual:

Field Source
Changelog / Commits Automated via git-cliff from conventional commits
Database migration detection Automated by CI scanning migration directories
Migration risk classification Automated by analysis tooling (Squawk)
Version bump recommendation Automated based on migration type + commit prefixes
API changes Manual: team documents new endpoints and breaking changes
Service dependencies Manual: team documents cross-service requirements
Rollback strategy narrative Manual: team assesses based on migration risk and data impact

Database Migration Policy

The Forward-Only Principle

Database rollbacks are unreliable in production for three fundamental reasons:

Partial failures leave unknown state. If an “up” migration adds two columns and fails after the first, the database is in a state the “down” migration doesn’t expect. Running the down migration will fail, leaving you stuck.

Reversing additive changes destroys data. Dropping a column that was successfully added doesn’t “undo” the addition. It permanently deletes all data in that column. Re-running the migration won’t restore it.

Previous code doesn’t contain future down migrations. When you roll back to a previous application version, that version’s codebase doesn’t include the down migrations for the changes that came after it. The rollback artifacts simply don’t exist in the image you’re deploying.

For these reasons, we treat database migrations as forward-only operations and use the Expand-Contract pattern for breaking schema changes.

Migration Risk Classifications

The CI pipeline classifies every detected migration and recommends a version bump based on the Version Bump Decision Guide above.

The Expand-Contract Pattern

Any schema change that would break the currently-running application code requires the Expand-Contract pattern. The core idea is to split one breaking change into a series of backward-compatible steps, each deployed as its own release. Because each step is backward-compatible, it can be rolled back independently without data loss (until the final cleanup step).

The pattern works in five steps:

Step 1: Expand. Add the new structure (columns, tables) alongside the old. Update application code to write to both old and new structures, but continue reading from the old. At this point the new structure exists but is not yet trusted. This step is a MINOR release — it’s purely additive. Rolling back means dropping the new structure, which is safe because nothing depends on it yet.

Step 2: Backfill. Migrate existing data from the old structure into the new. This runs as a background task (batched updates, throttled to avoid overwhelming the database). The application still reads from the old structure, so this step doesn’t affect correctness. Rolling back means clearing the new structure and restarting the backfill.

Step 3: Switch reads. Once the backfill is complete and data exists consistently in both structures, switch the application to read from the new structure while still writing to both. This is the key transition point. Rolling back means switching reads back to the old structure, which still has consistent data because dual-writes are still active.

Step 4: Stop writing to old structure. The old structure is now neither read from nor written to. This step is harder to roll back because the old structure stops receiving updates. However, there should be little need to revert since the new structure is already fully validated.

Step 5: Contract. Remove the old structure. This is the only truly irreversible step — it’s a MAJOR release. It should only happen after the new structure has been running in production long enough to be confident it’s correct.

How this maps to releases: Steps 1 through 4 are each a MINOR release (backward-compatible, additive or neutral). Step 5 is the only MAJOR release. If anything goes wrong at Steps 1-3, you can roll back that step’s release safely. The risk is concentrated in Step 5, which only runs after everything else is validated.

For implementation details including Rails and Python code examples, see the Database Migration Patterns Reference.

Migration Analysis in CI

Every merge to main triggers migration analysis that feeds into the draft release (pre-release):

  • Rails and Python services: Migration DSLs are converted to raw SQL, which is then linted by Squawk in CI.
  • Output: The analysis produces a risk classification, a recommended version bump, and any warnings (e.g., “this ALTER may lock the table for extended periods on tables over 1M rows”). This output is included in the draft release for the team to review.

The analysis is advisory, not blocking. It informs the team’s release decision rather than preventing merges or deploys.


Multi-Service Coordination

Terminology: Throughout this document, “service” refers to what we sometimes call an “application”. We use “service” here to emphasize the inter-service communication perspective.

Our architecture means each service maintains its own database and communicates with other services via APIs and message queues. Teams deploy independently on their own schedule. This works well because of our long-standing practice of maintaining backward compatibility across service boundaries.

Backward Compatibility Across Service Boundaries

We maintain backward compatibility on all inter-service contracts as our primary coordination mechanism. This applies to REST APIs, SNS/SQS message schemas, and any other interface between services. The key policies are:

New endpoints, new optional fields, and new message attributes are always safe. These are additive changes that don’t affect existing consumers. They map to MINOR releases.

Removing or renaming fields, changing response/message shapes, or adding required parameters are breaking changes. These require coordinated deployment across producer and consumer services, and they map to MAJOR releases on the producing service.

Deprecation before removal. When a field, endpoint, or message attribute needs to go away, mark it deprecated in one release, communicate the timeline to consuming teams, and remove it in a future release after consumers have migrated.

Coordinating Cross-Service Releases

When a release on one service requires a corresponding change on another, coordination relies on release documentation and team communication:

  1. The releasing team lists dependent services (and minimum compatible versions) under Service Dependencies in the release notes, making the requirement visible during review.
  2. Teams coordinate deployment ordering. Typically the backward-compatible side deploys first. For example: deploy the new API endpoint on Service A, then deploy the consumer code on Service B that uses it.

Configuration Coordination

Each service has its own repository containing application code and Dockerfile. Kubernetes deployment configuration (manifests, Kustomize overlays, environment-specific settings) lives separately in the deployment-config repository.

Because application code and deployment configuration are released independently, a complete deployment is the combination of a service image version and a deployment configuration version.

When a release requires a corresponding configuration change (e.g., new environment variables, updated resource limits), both should be documented in the release notes and coordinated during promotion.


Hotfix and Emergency Releases

When Production is broken and a fix is urgent, the standard release flow still applies, but with a compressed timeline. The goal is to keep the process intentional even under pressure, because emergencies are exactly when undocumented, unreviewed changes cause the most damage.

The key policies for emergency releases:

  • The fix must land on main before the hotfix branch is closed. This guarantees the fix is part of mainline history and will not be lost in future releases. If main is close to the broken tag, fix main first and cherry-pick to a hotfix branch. If main has diverged significantly, fix the hotfix branch first, deploy, close the incident — then port to main immediately after. Do not let conflict resolution on main delay a production fix.
  • Hotfix releases should always be PATCH releases. If the fix itself requires a database migration or a breaking change, that is a sign the underlying issue needs a more considered approach rather than an emergency patch.
  • Staging validation can be expedited but should not be skipped. Deploying an untagged commit directly to Production undermines traceability, which is exactly what you need most during an incident.
  • Tagging, release documentation, and Staging promotion are never skipped. Even under time pressure, these are what make the fix traceable and auditable.

For the step-by-step hotfix workflow (branching, cherry-picking, expedited promotion), see Hotfix and Emergency Releases in the Workflow Guide.


What qualifies as a hotfix

Before compressing the release timeline, the first question is whether the situation actually warrants it. The standard release process exists for a reason — skipping steps under false urgency is how undocumented changes accumulate.

A hotfix is warranted when something is broken in production right now and cannot wait until the next business day without meaningful user or business impact. The bar is intentionally high.

    flowchart TD
    A([After-hours deploy needed?]) --> B{Is something broken\nin production right now?}

    B -- No --> C[Not a hotfix.\nWait for business hours\nor the normal release process.]

    B -- Yes --> D{Can it wait\nuntil tomorrow morning?}

    D -- Yes --> E[It waits.\nSchedule for the next business day.]

    D -- No --> F{Manager aware\nand signed off?}

    F -- No --> G[Stop.\nContact your manager\nbefore proceeding.]

    F -- Yes --> H{Incident opened +\nrollback plan documented?}

    H -- No --> I[Open incident first.\nDocument rollback plan.\nThen deploy.]

    H -- Yes --> J([✅ Proceed with compressed release flow])

    style C fill:#f5f5f5,stroke:#ccc
    style E fill:#f5f5f5,stroke:#ccc
    style G fill:#fdecea,stroke:#e57373
    style I fill:#fff8e1,stroke:#ffb300
    style J fill:#e8f5e9,stroke:#66bb6a
  

Situations that do not qualify as hotfixes:

  • A feature that was recently shipped and has a minor gap or missing behaviour
  • Content availability issues affecting a small number of users on a low-traffic product
  • Anything that can reasonably wait until the next business day

When in doubt, ask your manager. A five-minute conversation before a deploy is always worth it.


If you are deploying after hours, it is an incident

Any after-hours production deploy — regardless of perceived risk — must be opened as an incident in the incident management tool before the deploy happens. This is not bureaucracy. It is the mechanism that creates visibility, traceability, and a rollback record. Without it, there is no paper trail if something goes wrong, and no way for other teams to know what changed.

Minimum bar for any after-hours deploy:

Requirement Why
Incident opened in incident management tool Creates a traceable record before any change is made
MR linked to a bug report or incident — not just the original feature Ensures the fix is independently trackable
One-line Slack message in the relevant channel before deploying No one should find out from the deployment itself
Manager aware and signed off No surprises. Awareness is the floor; sign-off is expected for non-trivial risk
Rollback plan documented, even in one sentence Forces the question before the deploy, not after

Shared systems

For production systems with multiple teams deploying to them, an undocumented after-hours deploy by one team creates risk for all of them. The policies above apply regardless of which team owns the change. If your deploy touches a shared system, the manager awareness requirement extends to the leads of other affected teams.


Success Metrics

These metrics help us evaluate whether the transition to intentional releases is achieving its goals:

Release documentation coverage: Do releases consistently include the required metadata (changelog, migration risk, dependencies, rollback strategy)? Measured by CI validation on release creation.

Migration success rate: How often do database migrations succeed without requiring manual intervention? Measured by tracking migration failures in Staging and Production.

Deployment visibility: Can any engineer explain what changed in a given deployment by reading the release documentation? Validated through periodic team check-ins.

Rollback clarity: Before Production deployments, can the team articulate the rollback strategy and its limitations? This is part of the approval step.

Mean time to recovery (MTTR): Does the new process improve our ability to recover from deployment issues, either through faster rollbacks (for PATCH releases) or through better-informed roll-forward decisions (for MINOR/MAJOR releases)?