Skip to content

Day 1 Homework and Wrap-Up

This pulls together everything from 1.2-testing-coverage.md, 1.3-dbt-mesh.md, and 1.4-dbt-mesh-part-2.md into one closing checklist, plus two new requirements: a documentation/test-coverage health check, and a final PR.

Work on your dpg/<your_name>/day1 branch. Keep committing as you go and don't save it all for one giant commit at the end; the PR review is easier with a real commit history.


1. Finish the outstanding exercise steps

From 1.2-testing-coverage.md

  • Use dbt Catalog's Recommendations (Testing category) to get a prioritized coverage-gap list
  • Write the tests yourself to fill the identified gaps
  • Run dbt build and confirm those specific gaps are closed
  • Identify which of the four accepted_values tests on the CRM data is most at risk of breaking from a system update (check the README for context on CRM volatility)
  • Apply the appropriate fix for that test - padding the accepted list, fixing casing/inconsistency in the staging SQL, or adding a severity - and justify why that tactic over the others in your commit message
  • Write the membership-fee test: basic=€599, premium=€1499, standard=€999, with a 12.5% discount applied for anyone in staff_members.csv
  • Update _crm__sources.yml documentation using the README as source material, writing it by hand (a good real-world candidate for dbt Wizard - done manually for this training)
  • Override the built-in not_null test with an optional null_values argument, so specific strings (e.g. "N/A") can be treated as null for a given test without changing default behavior everywhere else
  • Apply the overridden test to a real column with this problem, and confirm an unmodified not_null test elsewhere in the project still behaves exactly as before
  • Extension 2: make the null_values leniency configurable at the project level instead of hardcoded per test - add a boolean var (e.g. na_is_null) under vars: in dbt_project.yml, and reference it in a where config on the test (e.g. where: "{{ 'true' if var('na_is_null', true) else 'false' }}") so the leniency can be switched off project-wide - in CI, or once the upstream "N/A" cleanup is actually finished - without touching every test's YAML
  • Tag a handful of your most critical tests (e.g. important) and confirm dbt build --select tag:important runs only those
  • Add store_failures: true to a not_null test on one of fct_streaming_events's join-enriched columns (ctnt_type, release_date, or runtime_minutes), and confirm the failures table shows what the null rows actually have in common

From 1.3-dbt-mesh.md

  • Write the cross-project analyses/ query against the base project's streaming model
  • Answer all four investigation questions (which model, grain comparison, legacy/new overlap window, why fct/dim models aren't available to you)
  • Build the intermediate model reshaping legacy streaming data to match the new data's shape
  • Build fct_all_streaming_events, consolidating legacy + new, with an is_legacy flag and SQL comments explaining any columns that are empty on one side
  • Write the yml for fct_all_streaming_events yourself, including tests that only apply to legacy or new rows via a where filter
  • Extension 1: add versioning to the models involved in this cross-project dependency, and write down what happens if base updates a model you're depending on while you have a version pinned

From 1.4-dbt-mesh-part-2.md

  • Add an enforced contract to fct_crm_touchpoints, with every column's data_type verified against the actual warehouse output (not guessed)
  • Confirm fct_crm_touchpoints is materialized as table or view, not ephemeral
  • Set access: public on fct_crm_touchpoints and on the streaming domain's front-door model; set access: private on everything upstream of those (dim_sales_reps, dim_advertisers, fct_advertiser_contracts, and the internal streaming models)
  • Create two groups (CRM, streaming), each with a named owner, and assign every model in both domains so nothing is left ungrouped
  • Add the v1/v2 versioning block to fct_crm_touchpoints splitting occurred_at into occurred_at_date and occurred_at_time, with latest_version: 1 and a deprecation_date on v1
  • Confirm dbt run --select fct_crm_touchpoints builds both versions, and dbt run --select fct_crm_touchpoints,version:latest builds only v1

2. Project health check - documentation & test coverage

Don't aim for 100% coverage on either front - that's a vanity metric, not a quality bar. Instead, confirm the following are all true before moving on:

Documentation:

  • Every model has a top-level description
  • Every source table has a top-level description (this should already be true from the _crm__sources.yml update above)
  • Any column whose meaning isn't obvious from its name has a description - timestamps, status/enum fields, anything derived or calculated
  • dbt docs generate runs clean with no missing-doc warnings on the models you touched today

Test coverage:

  • Every primary key has not_null + unique
  • Every foreign key has not_null, and a relationships test where the parent table is trustworthy enough to assert against
  • Every enum-like column (status fields, type fields) has an accepted_values test, with the severity/tolerance decided in section 1 applied
  • The membership-fee logic from section 1 has a real test, not just a manual spot-check
  • dbt build runs clean - zero failing tests - across everything you touched today

If a fresh look at Catalog's coverage report still shows gaps after this, use judgment: a gap on a business-critical column is worth fixing, a gap on a free-text notes column usually isn't. Write one line in your PR description on any gap you deliberately chose to leave open and why.


3. Open a PR

Open a PR from your dpg/<your_name>/day1 branch. The description should cover, at minimum:

  • Summary - one or two sentences on what today's work accomplished overall
  • Testing coverage - what gaps existed, what you closed, and what (if anything) you deliberately left open and why
  • CRM test parameters - which accepted_values test you changed, which tactic you picked, and why that tactic over the alternatives
  • New test + docs - a note on the membership-fee test and the _crm__sources.yml update
  • Custom test + tagging - a note on the not_null override, whether you wired up the Extension 2 project-variable toggle, and what store_failures showed you about the enriched-column nulls
  • dbt Mesh (Part 1) - what you learned querying the base project's streaming data, and a link to your answers on the investigation questions
  • dbt Mesh (Part 2 - governance) - confirm contracts, access, and groups are in place, and explicitly flag the breaking change: occurred_at is being split and will eventually disappear - call out the deprecation date so a reviewer (standing in for the consuming team) can't miss it
  • Versioning - a one-line note on what happens if base updates a model you have version-pinned, from the Extension 1 exercise

Tag the PR for review as if the "other reporting team" from 1.4 were genuinely about to depend on this - the description should give them everything they'd need without having to ask you directly.