Advanced Testing¶
This is the deepest testing content in the whole workshop. By now you've covered generic tests, singular tests, and dbt Mesh. This topic asks you to combine all three: configure a test deliberately (not just add it), and reason about where tests should live once a project boundary is involved.
Exercise¶
mediapulse_analytics's dim_legacy_content has an is_mapped column that's false by design for any StreamView catalog item that hasn't been migrated to a current MediaPulse content_id yet. A blanket not_null-style expectation that "everything should be mapped" would be wrong - the migration is deliberately incomplete. But you still want visibility into how incomplete it is.
Step 1 - Pick your signal¶
- Step complete
Decide how you'd monitor is_mapped coverage without treating an unmapped row as a hard failure. Consider: a warn-severity test versus a passing/failing threshold on the proportion of unmapped rows.
Step 2 - Apply severity and error_if/warn_if deliberately¶
- Step complete
Configure a test on is_mapped (or write a singular test if a generic test can't express your threshold) using severity, and error_if/warn_if if relevant, so that some level of unmapped rows is expected and tolerated, but a regression beyond your threshold still surfaces.
Hint: severity, error_if, warn_if
When severity is error (the default), dbt checks error_if first and falls back to warn_if; when severity is warn, it skips straight to warn_if. Both accept a comparison expression, not just a boolean. See severity, error_if, and warn_if for the exact config shape and how the two interact.
Step 3 - Scope it with where¶
- Step complete
If your threshold should only apply to a subset of rows (for example, content mapped before a cutoff date should have a much higher bar than content added last week), scope the test with a where config rather than filtering in the model itself.
Hint: where config on data tests
A where config filters which rows a test evaluates, without changing the model's SQL. It's documented alongside the other test configs in data test configurations.
Deliverable: the test definition (generic with config, or singular), and a short justification for the severity/threshold you chose given that the migration is incomplete on purpose.
Extension¶
Imagine a third project joins this mesh - say, a mediapulse_finance project that wants to ref() public models from mediapulse_base and mediapulse_analytics for revenue reporting. Not every test you've written so far should mean the same thing to that new consumer.
Step 1 - Classify your existing tests¶
- Step complete
Look back over the tests you and your group have added across mediapulse_base and mediapulse_analytics this workshop. For each one (or a representative sample), classify it as either a "public contract" - something any downstream consumer should be able to rely on without reading the model's SQL - or "internal-only" - a test that only matters to people actively developing inside that project.
Step 2 - Connect this to access and contracts¶
- Step complete
mediapulse_base already marks some models public and others protected/private in dbt_project.yml, which controls whether mediapulse_analytics can ref() them at all. Explain how a model's access level and its test coverage relate: should a model need a minimum level of test coverage before it's allowed to be marked public? Who would enforce that, and how?
Hint: access and governance are related but separate features
Access (public/protected/private) controls whether a ref() is even allowed to resolve across a project boundary; it says nothing on its own about test coverage. Model contracts are a separate, related governance feature that enforces a model's column names and data types at build time. Read Model governance for how access, contracts, and versioning fit together as a set of related controls, then decide which one(s) you'd actually use to enforce a testing bar.
Step 3 - Upgrade one test to public-contract grade¶
- Step complete
From your classification, pick one existing test that you marked (or should have marked) "public contract" but that isn't actually configured strongly enough to deserve that label yet - for example, a bare not_null on a foreign key that should really be a relationships test, or a generic test with no severity/where config guarding a column that genuinely needs one. Edit the real YAML in mediapulse_base or mediapulse_analytics to bring that one test up to the standard your policy describes.
Step 4 - Write the policy¶
- Step complete
Produce a short written policy (a few bullet points) a new team joining this mesh could follow: what must be true (test-wise) about a model before it can be marked public, and what's left to each project's own discretion.
Deliverable: your test classification from step 1, the upgraded test's YAML diff from step 3, and the written policy from step 4.
Done?
You've configured a test deliberately rather than reflexively, upgraded a real test to public-contract grade, and written a policy a new team could actually follow.
Now head to Dynamic data masking to close a governance gap access levels alone don't solve.