Skip to content
Intentional Release Workflow Guide

Intentional Release Workflow Guide

This document describes the concrete workflow for how we build, release, and deploy software. It is the practical companion to the Intentional Release Guidelines, which covers the principles and reasoning behind these decisions. For detailed database migration patterns (Expand-Contract implementations, code examples), see the Database Migration Patterns Reference.


Environment Strategy

Note: During the transition to GitOps, some services may still deploy via the old CI/CD pipelines. The descriptions below reflect the target state.

We use three environments with distinct purposes. Each environment answers a different question.

Dev answers: “Does this application build, deploy, and start correctly?” Dev streams the main branch continuously. Every merge to main triggers a build, and the resulting image is automatically deployed to Dev. Each application runs in isolation here — there are no connections to other applications in the Dev environment. Dev catches deployment and startup issues early but is not a quality gate for releases.

Staging answers: “Is this release safe and ready for production?” Staging only receives tagged releases. When a team creates a release tag (e.g., v2.1.0), the resulting image becomes available for promotion to Staging. Kargo AnalysisTemplates validate the release before it can proceed to Production. Staging runs with coordinated seed data across all services so cross-service integration can be verified meaningfully.

Production answers: “Is this running reliably for our users?” Production receives releases that have been validated in Staging. Promotion from Staging to Production requires manual approval (recommended) or can be configured for automatic promotion after Staging validation passes. The same image that passed Staging validation is what runs in Production. No rebuilds, no surprises.

The flow is: main branch deploys continuously to Dev. Release tags promote automatically to Staging for validation. After validation passes, a team member approves the Production promotion (manual approval is recommended, though automatic promotion after successful validation is also supported).

Why Dev Exists

Without Dev, we face a choice: either Staging streams main continuously (which makes it useless as a quality gate), or developers have no environment to verify their code deploys and runs correctly after merge. Dev gives teams that continuous feedback loop while keeping Staging clean for intentional release validation.

Dev might use the same seed data as Staging. It’s more practical than crafting a separate dataset or using random data, and it keeps environments consistent so issues caught in Dev are reproducible in Staging. Since Dev streams main continuously (including potentially broken migrations), Dev databases may be reset more frequently than Staging. Seed data is reapplied automatically on reset.


The Developer Workflow

Day-to-Day Development

Nothing changes about how developers write code. You still work on feature branches, open merge requests, get reviews, and merge to main. The change is what happens after merge.

Before (old way): Merge to main. CI builds an image. Image goes straight to staging or production. Hope nothing breaks.

After (new way): Merge to main. CI builds an image and runs pre-release analysis. Image deploys to Dev automatically. CI generates a draft release with changelog, migration analysis, and version bump recommendation. When ready, the team reviews the draft and creates a release tag. The tagged image promotes through Staging and Production with validation gates.

The crucial difference: deploying to production is now a conscious decision, not an automatic consequence of merging code.

Creating a Release

When a team decides it’s time to release:

  1. Review the draft release. The CI pipeline generates a draft release on every merge to main. This draft includes a changelog generated by git-cliff from conventional commits, database migration analysis with risk classification, a recommended version bump (PATCH, MINOR, or MAJOR) based on the analysis, and a list of commits since the last release.
  2. Verify the recommendation. Check that the version bump makes sense. Refer to the Version Bump Decision Guide in the Guidelines for the full decision tree.
  3. Create the tag. Accept or adjust the recommended version, then create the tag. This triggers the release pipeline, which generates the release notes automatically.
  4. Complete the release documentation. The pipeline-generated release notes cover changelog and migration analysis automatically. The releasing team must then update the notes to also document: API changes (new endpoints, breaking changes), service dependencies (which versions of other services are required), and the rollback strategy narrative (what happens if you need to roll back, including data impact). See the Release Documentation Standards for the full template and the breakdown of automated vs. manual fields.
  5. Monitor promotion. The tagged image flows through Staging validation automatically. If AnalysisTemplates pass, it becomes eligible for Production promotion. A team member reviews the release notes and migration risk, then approves the Production promotion. (Automatic promotion after Staging validation is possible but manual approval is recommended.)
    %%{init: {'theme': 'base', 'themeVariables': { 'background': '#00000000', 'primaryColor': '#e8eef7', 'primaryTextColor': '#000', 'primaryBorderColor': '#4a6fa5', 'secondaryColor': '#e6f4ea', 'secondaryTextColor': '#000', 'secondaryBorderColor': '#28a745', 'tertiaryColor': '#fdf3e3', 'tertiaryTextColor': '#000', 'tertiaryBorderColor': '#d4a043', 'actorBkg': '#e8eef7', 'actorTextColor': '#000', 'actorBorder': '#4a6fa5', 'actorLineColor': '#4a6fa5', 'noteBkgColor': '#fdf3e3', 'noteTextColor': '#000', 'noteBorderColor': '#d4a043', 'signalColor': '#4a6fa5', 'signalTextColor': '#000', 'labelBoxBkgColor': '#e8eef7', 'labelTextColor': '#000', 'labelBoxBorderColor': '#4a6fa5', 'loopTextColor': '#000', 'activationBkgColor': '#e8eef7', 'activationBorderColor': '#4a6fa5' }}}%%
sequenceDiagram
    participant Dev as Developer
    participant CI as CI Pipeline
    participant Reg as Container Registry
    participant GL as GitLab Releases
    participant K as Kargo
    participant Argo as ArgoCD
    participant S as Staging
    participant P as Production

    rect rgba(74, 111, 165, 0.12)
    Note over Dev,Argo: Every merge to main
    Dev->>CI: Merge to main
    CI->>CI: Build, test
    CI->>Reg: Push image (commit SHA tag)
    CI->>K: Image available
    K->>Argo: Auto-promote to Dev
    CI->>CI: Diff main vs latest tag
    CI->>CI: Analyze migrations (Squawk)
    CI->>CI: Generate changelog (git-cliff)
    CI->>GL: Create/update draft release
    end

    rect rgba(212, 160, 67, 0.12)
    Note over Dev,GL: Release decision (manual)
    Dev->>GL: Review draft release
    Dev->>GL: Publish release, create tag v1.3.0
    end

    rect rgba(40, 167, 69, 0.12)
    Note over CI,P: Promotion pipeline
    GL->>CI: Tag triggers release build
    CI->>Reg: Push image (v1.3.0 tag)
    CI->>CI: Diff new tag vs previous tag
    CI->>CI: Final migration analysis
    CI->>K: Tagged image available
    K->>Argo: Promote to Staging
    Argo->>S: Deploy to Staging
    K->>K: Run AnalysisTemplates
    Note over K: Health checks, smoke tests,<br/>migration verification
    K-->>K: Validation passes
    Dev->>K: Manual approval
    K->>Argo: Promote to Production
    Argo->>P: Deploy to Production
    end
  

What the CI Pipeline Does

On Every Merge to Main

After the standard build and test stages, the pipeline runs three additional analysis stages. These run on every merge, not only when migration files change. A merge with no migrations still produces a draft release recommending a PATCH bump.

Analyze migrations. The pipeline diffs main against the latest release tag to identify new or modified migration files. For each detected migration, it extracts the SQL (Rails and Python migration DSLs are converted to raw SQL, which is then linted by Squawk for PostgreSQL analysis) and classifies the risk: whether the change is additive (new tables, new columns), breaking (drops, type changes, constraint changes), or requires expand-contract coordination. When no migrations are detected, the analysis reports “no database changes” and recommends PATCH.

Generate changelog. git-cliff generates the changelog from conventional commits since the last release tag. The pipeline enriches this with database migration entries categorized by risk level and flags any detected breaking changes.

Create draft release. The pipeline creates or updates a draft release in GitLab with the enriched changelog, the recommended version bump, and the migration risk summary. This draft accumulates changes across multiple merges. Each merge updates the same draft until the team decides to publish it as a release.

All analysis is advisory, not blocking. It informs the team’s release decision rather than preventing merges or deploys.

On Tag Creation

When a release tag is created, the pipeline runs a final analysis pass. This time it diffs the new tag against the previous tag, producing the definitive migration analysis and changelog for the release. This ensures the published release documentation reflects exactly what’s in the tag, not an intermediate state from the draft.


GitOps Architecture

Repository Structure

We use a two-repository pattern:

Application repositories contain source code, Dockerfiles, CI pipeline definitions, and migration files. Each service has its own repo (or lives in its section of the monorepo). CI pipelines in these repos handle building, testing, and creating releases.

The deployment-config repository contains Kustomize manifests organized with base configurations and environment-specific overlays. This repo is the source of truth for what should be running in each environment. ArgoCD watches this repo.

How ArgoCD and Kargo Work Together

The following diagram shows the full flow from a tagged release through to production deployment:

    flowchart LR
    subgraph APP["Application Repo"]
        Tag["Release tag<br/><i>v1.3.0</i>"]
    end

    subgraph CI_P["CI Pipeline"]
        Build["Build & push<br/>image v1.3.0"]
    end

    subgraph REG["Container Registry"]
        Img["myapp:v1.3.0"]
    end

    subgraph KARGO["Kargo"]
        WH["Warehouse<br/><i>watches registry</i>"]
        FR["Freight<br/><i>versioned artifact bundle</i>"]
        WH --> FR
    end

    subgraph DEPLOY["Deployment-Config Repo"]
        PB["production/myapp branch<br/><i>image: v1.3.0</i>"]
        SB["staging/myapp branch<br/><i>image: v1.3.0</i>"]
    end

    subgraph ARGO["ArgoCD"]
        AP["Watches production branch"]
        AS["Watches staging branch"]
    end

    subgraph PC["Production Cluster"]
        PN["Production workloads"]
    end

    subgraph SC["Staging Cluster"]
        SN["Staging workloads"]
    end

    Tag --> Build --> Img --> WH
    FR -->|"manual approve<br/>to Production"| PB
    FR -->|"auto-promote<br/>to Staging"| SB
    PB --> AP --> PN
    SB --> AS --> SN

    style Tag fill:#ffc107,stroke:#d4a043,color:#000
    style Build fill:#e8eef7,stroke:#4a6fa5,color:#000
    style Img fill:#e8eef7,stroke:#4a6fa5,color:#000
    style WH fill:#e8eef7,stroke:#4a6fa5,color:#000
    style FR fill:#e8eef7,stroke:#4a6fa5,color:#000
    style SB fill:#e6f4ea,stroke:#28a745,color:#000
    style PB fill:#e6f4ea,stroke:#28a745,color:#000
    style AS fill:#e6f4ea,stroke:#28a745,color:#000
    style AP fill:#e6f4ea,stroke:#28a745,color:#000
    style SN fill:#e6f4ea,stroke:#28a745,color:#000
    style PN fill:#e6f4ea,stroke:#28a745,color:#000
  

Kargo manages the promotion pipeline. It watches for new artifacts (container images from tagged releases) via Warehouses, bundles them into Freight (versioned artifact sets), and promotes them through environments according to defined policies. Kargo uses a branch-per-app-per-environment pattern: when a promotion happens, Kargo creates or updates an environment-specific branch in the deployment-config repo with the new image tag.

ArgoCD watches those environment-specific branches and reconciles the cluster state. When Kargo updates the staging branch for a service, ArgoCD detects the change and deploys the new version to the staging cluster.

The full promotion flow:

  1. CI builds an image from a release tag and pushes it to the container registry.
  2. Kargo’s Warehouse detects the new image and creates a Freight.
  3. Kargo auto-promotes to Staging by updating the staging branch in the deployment-config repo.
  4. ArgoCD syncs the staging cluster to match.
  5. Kargo AnalysisTemplates run validation against the deployed release.
  6. On manual approval, Kargo promotes to Production by updating the production branch.
  7. ArgoCD syncs the production cluster.

Dev Environment Flow

Dev works differently from Staging and Production. Since Dev streams main continuously, CI pushes images tagged with the commit SHA. Kargo detects the new image and auto-promotes to Dev. ArgoCD syncs the dev cluster. This gives teams immediate feedback on whether their code deploys and runs correctly without waiting for a release. Kargo can still run an AnalysisTemplate after promotion to confirm deployment health, but since Dev is a terminal stage with no downstream environment, the result is informational only and won’t gate anything.


Validation Gates

Kargo AnalysisTemplates define what “validated in Staging” means. These are CRDs borrowed from the Argo Rollouts project and run automatically after a release is promoted to Staging.

AnalysisTemplates support multiple validation approaches: running containerized processes as Kubernetes Jobs (exit code 0 = success), making HTTP requests and evaluating JSON responses, and querying monitoring tools like Prometheus and Datadog. For the full capabilities, see the Kargo AnalysisTemplate reference.

What We Can Validate

Here are some examples of what AnalysisTemplates can check after a promotion:

Health checks. Basic readiness and liveness probe verification. The service starts, responds to health endpoints, and stays healthy for a defined observation period. For example, an AnalysisTemplate can use the web metric provider to hit the service’s /healthz endpoint every 30 seconds for X iterations, requiring a 2xx response each time. Up to Y failures are tolerated before the analysis fails. This catches services that start successfully but crash or degrade shortly after deployment.

Database migration safety. Verify that migrations applied cleanly and the service operates correctly against the new schema. This catches issues like missing indexes, constraint violations on existing data, or migrations that work on empty databases but fail with real seed data. Implemented as a Job-based AnalysisTemplate that runs migration verification scripts.

What Passes vs. Fails

A release passes validation when all configured AnalysisTemplates succeed within their timeout period. A failure blocks promotion to Production and alerts the team. The team can then investigate in Staging, fix forward with a new release, or roll back Staging to the previous release.


Seed Data Strategy

Meaningful Staging validation requires coordinated seed data across all services. Without it, cross-service tests break because references point to non-existent entities.

Shared Entity Registry

We maintain a shared registry of canonical test entity IDs. This is a simple configuration file that defines the well-known test entities all services should create: specific IDs with defined roles, related resources referencing known entities, and so on. The registry ensures referential integrity across service boundaries without coupling the services themselves.

How It Works

Each team is responsible for generating their own seed data following the shared registry. Each service creates the defined test entities with appropriate attributes, referencing the known IDs from the registry where cross-service relationships exist. Teams implement seed data generation independently. They know their own domain best.

A working group (to be established as part of Phase 3) defines the registry requirements and maintains the shared entity definitions. Validation scripts verify that all expected entities exist and cross-service references resolve correctly.

Seed data is required in Staging, where cross-service validation depends on consistent, well-known entities. Dev can also use seed data for convenience, but teams are free to use random or ad-hoc data there since nothing gates on it. When databases are reset (before testing a new release batch, on-demand, or after a broken migration), seed data is re-applied as part of the reset process.


Multi-Service Coordination

Independent Services

Most services release independently. They follow their own semantic versioning lifecycle, create tags when they’re ready, and promote through environments on their own schedule. This is the default and preferred mode, supported by maintaining backward compatibility across API boundaries.

When a field, endpoint, or message attribute needs to be removed, the producing service must deprecate it in one release, communicate the timeline to consuming teams, and remove it in a future release after consumers have migrated. See Backward Compatibility Across Service Boundaries in the Guidelines for the full policy.

Coordinated Releases

When services do need to deploy together (e.g., a backend API change that requires a matching frontend change), coordination is a team responsibility, not something the infrastructure handles. Each service has its own independent Kargo promotion pipeline.

Teams coordinate by listing dependent services and minimum compatible versions in the release documentation, and by promoting services to Staging in the correct order (backward-compatible side first). If a cross-service integration AnalysisTemplate is configured, Staging validation may catch version mismatches, but this is a safety net rather than a coordination mechanism.

Configuration Changes

Application configuration lives in the deployment-config repository, separate from application code. Configuration changes (environment variables, resource limits, etc.) follow the same promotion path through Kargo.

When a release requires a corresponding configuration change, the release notes should document the required config changes and the promotion should be coordinated.

If a configuration change needs to be promoted independently of an application release, a Kargo Freight can be manually created to target a specific commit in the deployment-config repository.


Rollback Strategy

Application Rollback

Rolling back the application code is straightforward with GitOps: promote the previous release tag to the environment. Kargo and ArgoCD handle the rest. This works for PATCH releases (no database changes) and is safe for MINOR releases where the old code is compatible with the new additive schema.

Database Rollback Reality

As documented in the Intentional Release Guidelines, database rollbacks in production are unreliable. The strategy depends on the version type:

PATCH releases (no DB changes): Roll back freely. The database hasn’t changed.

MINOR releases (additive DB changes): Rolling back the application code is safe because the old code ignores new columns and tables. The new database structures remain but are unused. Clean them up in a future release if needed, or leave them for the next attempt.

MAJOR releases (breaking DB changes): These use the Expand-Contract pattern, which means the “breaking” change is deployed as a series of backward-compatible releases. If something goes wrong at any phase, you can roll back that phase’s release safely: at the Expand phase, the new structures can simply be dropped since nothing depends on them yet; at the Switch Reads phase, reads can be switched back to the old structure since dual-writes keep it consistent. The final Contract step (removing old schema) is the only truly irreversible step, and it only happens after the new structure has been running in production long enough to be confident it’s correct. See the Intentional Release Guidelines for the full five-step breakdown.

Emergency Procedures

For critical production issues:

Application code issue: Promote the previous release tag immediately. Kargo handles the demotion.

Failed migration: Assess the database state and fix forward with a hotfix release. Do not attempt to run down migrations in production.

Data corruption: Restore from backup. This is a last resort and involves data loss for the recovery window.


Hotfix and Emergency Releases

When Production is broken and a fix is urgent, the standard release flow still applies, but with a compressed timeline. The goal is to keep the process intentional even under pressure.

The fix always lands on main first. Regardless of the release strategy, the fix is merged to main through the normal merge request process. This guarantees the fix is part of the mainline history and will not be lost in future releases.

When main Is Safe to Release

If the accumulated commits on main since the last release are low-risk (no unfinished features, no risky migrations), tag a new PATCH release from main HEAD and promote through Staging with expedited validation.

When main Has Accumulated Unreleased Changes

If main contains commits that are not yet ready for Production, tagging HEAD would ship unintended changes. In this case:

  1. Merge the fix to main through the normal merge request process. If main has diverged significantly and conflict resolution would delay the fix, create the hotfix branch first, deploy, close the incident, and merge to main immediately after.
  2. Create a hotfix branch from the broken tag. For example, if v1.3.0 is broken, branch from that tag: hotfix/v1.3.1.
  3. Cherry-pick the fix from main onto the hotfix branch.
  4. Tag the hotfix release (v1.3.1) from the hotfix branch.
  5. Promote through Staging with focused validation on the fix. Staging validation can be expedited but should not be skipped entirely.

Hotfix releases should always be PATCH releases. If the fix itself requires a database migration or a breaking change, that is a sign the underlying issue needs a more considered approach rather than an emergency patch.

What to Skip, and What Not To

Can be expedited: Staging validation can focus narrowly on the fix rather than full regression. Production approval can happen immediately after Staging passes.

Should not be skipped: Tagging, release documentation, and Staging promotion. Deploying an untagged commit directly to Production undermines traceability, which is exactly what you need most during an incident.

For the policy rationale behind these rules, see Hotfix and Emergency Releases in the Guidelines.