dbt Catalog¶
You're the new hire on mediapulse_base's ads, news, and podcasts domains. Before you touch any code today, use dbt Catalog to find out what's actually going on in the project you've inherited - lineage, test coverage, and a few rough edges you'll only fix properly in later topics.
Prerequisite: a production (or staging) deployment environment for mediapulse_base, with at least one successful dbt build job run, so Catalog has metadata to display.
Goal: get acquainted with your three domains and find the gaps worth flagging before you start writing SQL.
1. Scope the DAG with selectors¶
mediapulse_base also contains a streaming domain that isn't yours today. Before you go anywhere near the lineage graph, confirm your actual scope on the command line - you've used --select with a single model before, now you need it for three domains at once:
- What does
source:podcasts+select on its own, and what does the trailing+add thatsource:podcastsalone wouldn't? - Confirm nothing from
streamingshows up anywhere in the output.
Hint: combining selectors
Selectors separated by spaces are unioned - source:ads+ source:news+ source:podcasts+ returns every node matched by any of the three, not just ones matched by all of them (that's what a comma, with no space, would do instead). + after a node means "and everything downstream of it" (staging → intermediate → marts); + before a node means upstream instead. See Set operators and Graph operators for the rest.
2. Walk the lineage¶
Open the lineage graph for mediapulse_base - it should match the node list you just got on the command line - and trace ads, news, and podcasts from source through staging to marts.
- One staging model joins two sources at once, plus a seed. Find it, and note what each input contributes.
- The seed it uses is the only seed in this project. Find it and work out why it's there instead of a static list in the model itself.
- One raw source gets read twice, independently, on two separate paths through the DAG instead of once via its own staging model. Find it - this is exactly what Catalog's Source Fanout rule (see step 3) is warning you about.
Hint: streaming is not your problem today
You'll also see mediapulse_analytics depending on the streaming staging layer, plus one podcasts and one news staging model. That's real, but it's Group 2/3 territory once dbt Mesh is in play - note it and move on.
3. Read the Recommendations tab¶
Open Recommendations for mediapulse_base and filter to the Modeling category.
- Confirm the Multiple Sources Joined flag matches the staging model you found in step 2.
- Confirm the Source Fanout (or duplicate-source) flag matches the source you found being read twice.
Now filter to Testing. You should find a handful of Missing Primary Key Test flags - roughly half sit in podcasts/news. List which models they're on.
What Recommendations actually checks
Catalog's Recommendations are powered by dbt_project_evaluator and grouped into Modeling, Testing, Documentation, and Governance rules. See Project recommendations for the full rule list - you'll use more of these categories in later topics.
4. Find the dead ends¶
Still in the lineage graph: models/intermediate/ has more than one model with nothing downstream of it - built, tested, documented, and never ref()'d by a mart.
For each one you find, write one sentence on what it looks like it was meant to feed, and whether that's a wiring gap (an easy fix) or a sign the model doesn't have a home yet. You're not fixing these now - some get wired up for real later today and tomorrow.
5. Check source freshness¶
Look at the config: block on each source in news and podcasts - none currently define freshness or loaded_at_field. (ads doesn't even declare a source() - its staging models read the raw tables directly, so there's nothing to check there yet.)
In one sentence: why does that matter more for a source you don't control the load schedule of than for a mart you build yourself?
Hint: current freshness syntax
As of dbt 1.10, freshness and loaded_at_field are configured under config: at the source or table level, not as top-level source keys:
sources:
- name: news
config:
freshness:
warn_after: {count: 24, period: hour}
error_after: {count: 48, period: hour}
loaded_at_field: _etl_loaded_at
Deliverable: your selector command and what it confirmed about streaming (1), the multi-source-join model and the fanned-out source (2), your Missing Primary Key Test list (3), your dead-end intermediate models with a one-line plan each (4), and your freshness answer (5).
Done?
You've read a real project's lineage and test coverage through tooling instead of guessing from file names, and you've got a running list of gaps you'll close for real over the next few topics.