Intentional Release Workflow Guide
This document describes the concrete workflow for how we build, release, and deploy software. It is the practical companion to the Intentional Release Guidelines, which covers the principles and reasoning behind these decisions. For detailed database migration patterns (Expand-Contract implementations, code examples), see the Database Migration Patterns Reference.
Environment Strategy
Note: During the transition to GitOps, some services may still deploy via the old CI/CD pipelines. The descriptions below reflect the target state.
We use three environments with distinct purposes. Each environment answers a different question.
Dev answers: “Does this application build, deploy, and start correctly?” Dev streams the main branch continuously. Every merge to main triggers a build, and the resulting image is automatically deployed to Dev. Each application runs in isolation here — there are no connections to other applications in the Dev environment. Dev catches deployment and startup issues early but is not a quality gate for releases.
Staging answers: “Is this release safe and ready for production?” Staging only receives tagged releases. When a team creates a release tag (e.g., v2.1.0), the resulting image becomes available for promotion to Staging. Kargo AnalysisTemplates validate the release before it can proceed to Production. Staging runs with coordinated seed data across all services so cross-service integration can be verified meaningfully.
Production answers: “Is this running reliably for our users?” Production receives releases that have been validated in Staging. Promotion from Staging to Production requires manual approval (recommended) or can be configured for automatic promotion after Staging validation passes. The same image that passed Staging validation is what runs in Production. No rebuilds, no surprises.
The flow is: main branch deploys continuously to Dev. Release tags promote automatically to Staging for validation. After validation passes, a team member approves the Production promotion (manual approval is recommended, though automatic promotion after successful validation is also supported).
Why Dev Exists
Without Dev, we face a choice: either Staging streams main continuously (which makes it useless as a quality gate), or developers have no environment to verify their code deploys and runs correctly after merge. Dev gives teams that continuous feedback loop while keeping Staging clean for intentional release validation.
Dev might use the same seed data as Staging. It’s more practical than crafting a separate dataset or using random data, and it keeps environments consistent so issues caught in Dev are reproducible in Staging. Since Dev streams main continuously (including potentially broken migrations), Dev databases may be reset more frequently than Staging. Seed data is reapplied automatically on reset.
The Developer Workflow
Day-to-Day Development
Nothing changes about how developers write code. You still work on feature branches, open merge requests, get reviews, and merge to main. The change is what happens after merge.
Before (old way): Merge to main. CI builds an image. Image goes straight to staging or production. Hope nothing breaks.
After (new way): Merge to main. CI builds an image and runs pre-release analysis. Image deploys to Dev automatically. CI generates a draft release with changelog, migration analysis, and version bump recommendation. When ready, the team reviews the draft and creates a release tag. The tagged image promotes through Staging and Production with validation gates.
The crucial difference: deploying to production is now a conscious decision, not an automatic consequence of merging code.
Creating a Release
When a team decides it’s time to release:
- Review the draft release. The CI pipeline generates a draft release on every merge to
main. This draft includes a changelog generated by git-cliff from conventional commits, database migration analysis with risk classification, a recommended version bump (PATCH, MINOR, or MAJOR) based on the analysis, and a list of commits since the last release. - Verify the recommendation. Check that the version bump makes sense. Refer to the Version Bump Decision Guide in the Guidelines for the full decision tree.
- Create the tag. Accept or adjust the recommended version, then create the tag. This triggers the release pipeline, which generates the release notes automatically.
- Complete the release documentation. The pipeline-generated release notes cover changelog and migration analysis automatically. The releasing team must then update the notes to also document: API changes (new endpoints, breaking changes), service dependencies (which versions of other services are required), and the rollback strategy narrative (what happens if you need to roll back, including data impact). See the Release Documentation Standards for the full template and the breakdown of automated vs. manual fields.
- Monitor promotion. The tagged image flows through Staging validation automatically. If AnalysisTemplates pass, it becomes eligible for Production promotion. A team member reviews the release notes and migration risk, then approves the Production promotion. (Automatic promotion after Staging validation is possible but manual approval is recommended.)
%%{init: {'theme': 'base', 'themeVariables': { 'background': '#00000000', 'primaryColor': '#e8eef7', 'primaryTextColor': '#000', 'primaryBorderColor': '#4a6fa5', 'secondaryColor': '#e6f4ea', 'secondaryTextColor': '#000', 'secondaryBorderColor': '#28a745', 'tertiaryColor': '#fdf3e3', 'tertiaryTextColor': '#000', 'tertiaryBorderColor': '#d4a043', 'actorBkg': '#e8eef7', 'actorTextColor': '#000', 'actorBorder': '#4a6fa5', 'actorLineColor': '#4a6fa5', 'noteBkgColor': '#fdf3e3', 'noteTextColor': '#000', 'noteBorderColor': '#d4a043', 'signalColor': '#4a6fa5', 'signalTextColor': '#000', 'labelBoxBkgColor': '#e8eef7', 'labelTextColor': '#000', 'labelBoxBorderColor': '#4a6fa5', 'loopTextColor': '#000', 'activationBkgColor': '#e8eef7', 'activationBorderColor': '#4a6fa5' }}}%%
sequenceDiagram
participant Dev as Developer
participant CI as CI Pipeline
participant Reg as Container Registry
participant GL as GitLab Releases
participant K as Kargo
participant Argo as ArgoCD
participant S as Staging
participant P as Production
rect rgba(74, 111, 165, 0.12)
Note over Dev,Argo: Every merge to main
Dev->>CI: Merge to main
CI->>CI: Build, test
CI->>Reg: Push image (commit SHA tag)
CI->>K: Image available
K->>Argo: Auto-promote to Dev
CI->>CI: Diff main vs latest tag
CI->>CI: Analyze migrations (Squawk)
CI->>CI: Generate changelog (git-cliff)
CI->>GL: Create/update draft release
end
rect rgba(212, 160, 67, 0.12)
Note over Dev,GL: Release decision (manual)
Dev->>GL: Review draft release
Dev->>GL: Publish release, create tag v1.3.0
end
rect rgba(40, 167, 69, 0.12)
Note over CI,P: Promotion pipeline
GL->>CI: Tag triggers release build
CI->>Reg: Push image (v1.3.0 tag)
CI->>CI: Diff new tag vs previous tag
CI->>CI: Final migration analysis
CI->>K: Tagged image available
K->>Argo: Promote to Staging
Argo->>S: Deploy to Staging
K->>K: Run AnalysisTemplates
Note over K: Health checks, smoke tests,<br/>migration verification
K-->>K: Validation passes
Dev->>K: Manual approval
K->>Argo: Promote to Production
Argo->>P: Deploy to Production
end
What the CI Pipeline Does
On Every Merge to Main
After the standard build and test stages, the pipeline runs three additional analysis stages. These run on every merge, not only when migration files change. A merge with no migrations still produces a draft release recommending a PATCH bump.
Analyze migrations. The pipeline diffs main against the latest release tag to identify new or modified migration files. For each detected migration, it extracts the SQL (Rails and Python migration DSLs are converted to raw SQL, which is then linted by Squawk for PostgreSQL analysis) and classifies the risk: whether the change is additive (new tables, new columns), breaking (drops, type changes, constraint changes), or requires expand-contract coordination. When no migrations are detected, the analysis reports “no database changes” and recommends PATCH.
Generate changelog. git-cliff generates the changelog from conventional commits since the last release tag. The pipeline enriches this with database migration entries categorized by risk level and flags any detected breaking changes.
Create draft release. The pipeline creates or updates a draft release in GitLab with the enriched changelog, the recommended version bump, and the migration risk summary. This draft accumulates changes across multiple merges. Each merge updates the same draft until the team decides to publish it as a release.
All analysis is advisory, not blocking. It informs the team’s release decision rather than preventing merges or deploys.
On Tag Creation
When a release tag is created, the pipeline runs a final analysis pass. This time it diffs the new tag against the previous tag, producing the definitive migration analysis and changelog for the release. This ensures the published release documentation reflects exactly what’s in the tag, not an intermediate state from the draft.
GitOps Architecture
Repository Structure
We use a two-repository pattern:
Application repositories contain source code, Dockerfiles, CI pipeline definitions, and migration files. Each service has its own repo (or lives in its section of the monorepo). CI pipelines in these repos handle building, testing, and creating releases.
The deployment-config repository contains Kustomize manifests organized with base configurations and environment-specific overlays. This repo is the source of truth for what should be running in each environment. ArgoCD watches this repo.
How ArgoCD and Kargo Work Together
The following diagram shows the full flow from a tagged release through to production deployment:
flowchart LR
subgraph APP["Application Repo"]
Tag["Release tag<br/><i>v1.3.0</i>"]
end
subgraph CI_P["CI Pipeline"]
Build["Build & push<br/>image v1.3.0"]
end
subgraph REG["Container Registry"]
Img["myapp:v1.3.0"]
end
subgraph KARGO["Kargo"]
WH["Warehouse<br/><i>watches registry</i>"]
FR["Freight<br/><i>versioned artifact bundle</i>"]
WH --> FR
end
subgraph DEPLOY["Deployment-Config Repo"]
PB["production/myapp branch<br/><i>image: v1.3.0</i>"]
SB["staging/myapp branch<br/><i>image: v1.3.0</i>"]
end
subgraph ARGO["ArgoCD"]
AP["Watches production branch"]
AS["Watches staging branch"]
end
subgraph PC["Production Cluster"]
PN["Production workloads"]
end
subgraph SC["Staging Cluster"]
SN["Staging workloads"]
end
Tag --> Build --> Img --> WH
FR -->|"manual approve<br/>to Production"| PB
FR -->|"auto-promote<br/>to Staging"| SB
PB --> AP --> PN
SB --> AS --> SN
style Tag fill:#ffc107,stroke:#d4a043,color:#000
style Build fill:#e8eef7,stroke:#4a6fa5,color:#000
style Img fill:#e8eef7,stroke:#4a6fa5,color:#000
style WH fill:#e8eef7,stroke:#4a6fa5,color:#000
style FR fill:#e8eef7,stroke:#4a6fa5,color:#000
style SB fill:#e6f4ea,stroke:#28a745,color:#000
style PB fill:#e6f4ea,stroke:#28a745,color:#000
style AS fill:#e6f4ea,stroke:#28a745,color:#000
style AP fill:#e6f4ea,stroke:#28a745,color:#000
style SN fill:#e6f4ea,stroke:#28a745,color:#000
style PN fill:#e6f4ea,stroke:#28a745,color:#000
Kargo manages the promotion pipeline. It watches for new artifacts (container images from tagged releases) via Warehouses, bundles them into Freight (versioned artifact sets), and promotes them through environments according to defined policies. Kargo uses a branch-per-app-per-environment pattern: when a promotion happens, Kargo creates or updates an environment-specific branch in the deployment-config repo with the new image tag.
ArgoCD watches those environment-specific branches and reconciles the cluster state. When Kargo updates the staging branch for a service, ArgoCD detects the change and deploys the new version to the staging cluster.
The full promotion flow:
- CI builds an image from a release tag and pushes it to the container registry.
- Kargo’s Warehouse detects the new image and creates a Freight.
- Kargo auto-promotes to Staging by updating the staging branch in the deployment-config repo.
- ArgoCD syncs the staging cluster to match.
- Kargo AnalysisTemplates run validation against the deployed release.
- On manual approval, Kargo promotes to Production by updating the production branch.
- ArgoCD syncs the production cluster.
Dev Environment Flow
Dev works differently from Staging and Production. Since Dev streams main continuously, CI pushes images tagged with the commit SHA. Kargo detects the new image and auto-promotes to Dev. ArgoCD syncs the dev cluster. This gives teams immediate feedback on whether their code deploys and runs correctly without waiting for a release. Kargo can still run an AnalysisTemplate after promotion to confirm deployment health, but since Dev is a terminal stage with no downstream environment, the result is informational only and won’t gate anything.
Validation Gates
Kargo AnalysisTemplates define what “validated in Staging” means. These are CRDs borrowed from the Argo Rollouts project and run automatically after a release is promoted to Staging.
AnalysisTemplates support multiple validation approaches: running containerized processes as Kubernetes Jobs (exit code 0 = success), making HTTP requests and evaluating JSON responses, and querying monitoring tools like Prometheus and Datadog. For the full capabilities, see the Kargo AnalysisTemplate reference.
What We Can Validate
Here are some examples of what AnalysisTemplates can check after a promotion:
Health checks. Basic readiness and liveness probe verification. The service starts, responds to health endpoints, and stays healthy for a defined observation period. For example, an AnalysisTemplate can use the web metric provider to hit the service’s /healthz endpoint every 30 seconds for X iterations, requiring a 2xx response each time. Up to Y failures are tolerated before the analysis fails. This catches services that start successfully but crash or degrade shortly after deployment.
Database migration safety. Verify that migrations applied cleanly and the service operates correctly against the new schema. This catches issues like missing indexes, constraint violations on existing data, or migrations that work on empty databases but fail with real seed data. Implemented as a Job-based AnalysisTemplate that runs migration verification scripts.
What Passes vs. Fails
A release passes validation when all configured AnalysisTemplates succeed within their timeout period. A failure blocks promotion to Production and alerts the team. The team can then investigate in Staging, fix forward with a new release, or roll back Staging to the previous release.
Seed Data Strategy
Meaningful Staging validation requires coordinated seed data across all services. Without it, cross-service tests break because references point to non-existent entities.
Shared Entity Registry
We maintain a shared registry of canonical test entity IDs. This is a simple configuration file that defines the well-known test entities all services should create: specific IDs with defined roles, related resources referencing known entities, and so on. The registry ensures referential integrity across service boundaries without coupling the services themselves.
How It Works
Each team is responsible for generating their own seed data following the shared registry. Each service creates the defined test entities with appropriate attributes, referencing the known IDs from the registry where cross-service relationships exist. Teams implement seed data generation independently. They know their own domain best.
A working group (to be established as part of Phase 3) defines the registry requirements and maintains the shared entity definitions. Validation scripts verify that all expected entities exist and cross-service references resolve correctly.
Seed data is required in Staging, where cross-service validation depends on consistent, well-known entities. Dev can also use seed data for convenience, but teams are free to use random or ad-hoc data there since nothing gates on it. When databases are reset (before testing a new release batch, on-demand, or after a broken migration), seed data is re-applied as part of the reset process.
Multi-Service Coordination
Independent Services
Most services release independently. They follow their own semantic versioning lifecycle, create tags when they’re ready, and promote through environments on their own schedule. This is the default and preferred mode, supported by maintaining backward compatibility across API boundaries.
When a field, endpoint, or message attribute needs to be removed, the producing service must deprecate it in one release, communicate the timeline to consuming teams, and remove it in a future release after consumers have migrated. See Backward Compatibility Across Service Boundaries in the Guidelines for the full policy.
Coordinated Releases
When services do need to deploy together (e.g., a backend API change that requires a matching frontend change), coordination is a team responsibility, not something the infrastructure handles. Each service has its own independent Kargo promotion pipeline.
Teams coordinate by listing dependent services and minimum compatible versions in the release documentation, and by promoting services to Staging in the correct order (backward-compatible side first). If a cross-service integration AnalysisTemplate is configured, Staging validation may catch version mismatches, but this is a safety net rather than a coordination mechanism.
Configuration Changes
Application configuration lives in the deployment-config repository, separate from application code. Configuration changes (environment variables, resource limits, etc.) follow the same promotion path through Kargo.
When a release requires a corresponding configuration change, the release notes should document the required config changes and the promotion should be coordinated.
If a configuration change needs to be promoted independently of an application release, a Kargo Freight can be manually created to target a specific commit in the deployment-config repository.
Rollback Strategy
Application Rollback
Rolling back the application code is straightforward with GitOps: promote the previous release tag to the environment. Kargo and ArgoCD handle the rest. This works for PATCH releases (no database changes) and is safe for MINOR releases where the old code is compatible with the new additive schema.
Database Rollback Reality
As documented in the Intentional Release Guidelines, database rollbacks in production are unreliable. The strategy depends on the version type:
PATCH releases (no DB changes): Roll back freely. The database hasn’t changed.
MINOR releases (additive DB changes): Rolling back the application code is safe because the old code ignores new columns and tables. The new database structures remain but are unused. Clean them up in a future release if needed, or leave them for the next attempt.
MAJOR releases (breaking DB changes): These use the Expand-Contract pattern, which means the “breaking” change is deployed as a series of backward-compatible releases. If something goes wrong at any phase, you can roll back that phase’s release safely: at the Expand phase, the new structures can simply be dropped since nothing depends on them yet; at the Switch Reads phase, reads can be switched back to the old structure since dual-writes keep it consistent. The final Contract step (removing old schema) is the only truly irreversible step, and it only happens after the new structure has been running in production long enough to be confident it’s correct. See the Intentional Release Guidelines for the full five-step breakdown.
Emergency Procedures
For critical production issues:
Application code issue: Promote the previous release tag immediately. Kargo handles the demotion.
Failed migration: Assess the database state and fix forward with a hotfix release. Do not attempt to run down migrations in production.
Data corruption: Restore from backup. This is a last resort and involves data loss for the recovery window.
Hotfix and Emergency Releases
When Production is broken and a fix is urgent, the standard release flow still applies, but with a compressed timeline. The goal is to keep the process intentional even under pressure.
The fix always lands on main first. Regardless of the release strategy, the fix is merged to main through the normal merge request process. This guarantees the fix is part of the mainline history and will not be lost in future releases.
When main Is Safe to Release
If the accumulated commits on main since the last release are low-risk (no unfinished features, no risky migrations), tag a new PATCH release from main HEAD and promote through Staging with expedited validation.
When main Has Accumulated Unreleased Changes
If main contains commits that are not yet ready for Production, tagging HEAD would ship unintended changes. In this case:
- Merge the fix to
mainthrough the normal merge request process. Ifmainhas diverged significantly and conflict resolution would delay the fix, create the hotfix branch first, deploy, close the incident, and merge tomainimmediately after. - Create a hotfix branch from the broken tag. For example, if
v1.3.0is broken, branch from that tag:hotfix/v1.3.1. - Cherry-pick the fix from
mainonto the hotfix branch. - Tag the hotfix release (
v1.3.1) from the hotfix branch. - Promote through Staging with focused validation on the fix. Staging validation can be expedited but should not be skipped entirely.
Hotfix releases should always be PATCH releases. If the fix itself requires a database migration or a breaking change, that is a sign the underlying issue needs a more considered approach rather than an emergency patch.
What to Skip, and What Not To
Can be expedited: Staging validation can focus narrowly on the fix rather than full regression. Production approval can happen immediately after Staging passes.
Should not be skipped: Tagging, release documentation, and Staging promotion. Deploying an untagged commit directly to Production undermines traceability, which is exactly what you need most during an incident.
For the policy rationale behind these rules, see Hotfix and Emergency Releases in the Guidelines.