Deployment & CI/CD¶
Group 2 - Part 2
Everything you've built so far has been validated by you, manually, in development. This topic is about what happens when someone opens a pull request instead - and specifically, what changes about that picture once there are two dbt projects in a mesh instead of one.
Exercise¶
Step 1 - Understand the default CI behaviour¶
- Step complete
Read up on what a dbt Cloud CI job actually runs by default, and what it needs in order to work.
Hint: The default selector
A CI job's default build command is dbt build --select state:modified+ - "everything that changed, plus everything downstream of it." Read Continuous integration jobs in dbt for how this is configured and what "downstream" means in practice.
- Which command does this rely on?
- What does "deferral" mean here, and which environment does dbt compare against by default?
Step 2 - Trace a one-project change¶
- Step complete
Imagine a PR that only edits a column description in stg_ads__campaigns in mediapulse_base. Would state:modified+ include any other models in that build? Would it include fct_ad_impressions? Justify your answer using what state:modified actually detects as a "modification."
Step 3 - Trace a cross-project change¶
- Step complete
Now imagine a PR that changes the logic inside stg_news__articles in mediapulse_base - a model that mediapulse_analytics's fct_content_performance depends on via dependencies.yml. If you only run CI on mediapulse_base's PR, does mediapulse_analytics get rebuilt or re-tested at all? What would have to be true for it to?
Hint: Cross-project refs resolve to an environment, not a manifest
In a dbt Mesh setup, ref() calls that cross a project boundary resolve to a specific environment (staging or production) of the upstream project, not to a snapshot bundled into the downstream project's own state. Read About dbt Mesh and Deployment environments for how that resolution works, and think about what that implies for whether a platform-only CI run is enough on its own.
Step 4 - Name the gap¶
- Step complete
Based on what you found in step 3, write down - in your own words - what a two-project mesh needs from CI that a single project doesn't: specifically, what has to happen for a change in mediapulse_base to be caught by mediapulse_analytics's own tests before it reaches production.
Deliverable: Your answers to steps 2-4, written as if explaining to a teammate why "just run CI on the project I changed" isn't automatically sufficient once a mesh dependency exists.
Extension¶
You've identified the gap. Now design around it, and this time put it into a real file dbt can actually use.
Step 1 - Draft a job plan for both projects¶
- Step complete
List the CI/CD jobs you'd want to exist across mediapulse_base and mediapulse_analytics together (for example: a PR-triggered job per project, plus anything that needs to run in response to the other project's changes). For each job, state its trigger and its selector logic in plain English.
Step 2 - Turn at least two jobs into a real selectors.yml¶
- Step complete
Create a selectors.yml file at the root of mediapulse_analytics and define at least two named selectors from your job plan - for example, one matching your "PR to this project" job and one matching your "something changed upstream in mediapulse_base" job. Give each a name and a description, and build its definition from real selection criteria (graph operators, state:modified+, and so on) rather than a placeholder.
Hint: selectors.yml shape and where it lives
A selectors.yml file sits at your project's root, next to dbt_project.yml, and holds a top-level selectors: list. Each entry needs a name and a definition - the definition can be a simple CLI-style string (like state:modified+) or a fuller method/value object if you need more control. See YAML selectors for the exact structure and how to reference a selector later with --selector <name>.
Step 3 - Decide what "deferred environment" each project needs¶
- Step complete
Read Defer and work out: does mediapulse_analytics's CI job need its own production environment to defer to, mediapulse_base's, or both? What breaks if one of them doesn't exist yet?
Step 4 - Handle the ordering problem¶
- Step complete
If a single commit touches both projects (unlikely in this workshop's setup, but common in real monorepo-style mesh setups), which project's job should run first, and does the second job need to wait for the first to finish and publish new state before it can defer correctly?
Step 5 - Write your recommendation¶
- Step complete
Summarise, in a short paragraph, the minimum CI setup you'd tell a new team to put in place the day they split a single dbt project into a two-project mesh like this one - not the ideal end state, the minimum that avoids a platform change silently breaking a downstream analytics project.
Deliverable: Your job plan (step 1), the selectors.yml file from step 2, and your minimum-viable recommendation (step 5).
Done?
You've designed and partly encoded a CI/CD plan that treats mediapulse_base and mediapulse_analytics as the connected projects they actually are, not two independent codebases.
That's Group 2 complete - well done. You've reviewed Catalog, gone deeper on testing, fixed a real macro/schema gap, and worked dbt Mesh, project evaluation, masking, and CI/CD across both projects.