Skip to content

dbt Catalog

You're the new hire on mediapulse_base's ads, news, and podcasts domains. Before you touch any code today, use dbt Catalog to find out what's actually going on in the project you've inherited - lineage, test coverage, and a few rough edges you'll only fix properly in later topics.

Prerequisite: a production (or staging) deployment environment for mediapulse_base, with at least one successful dbt build job run, so Catalog has metadata to display.

Goal: get acquainted with your three domains and find the gaps worth flagging before you start writing SQL.

1. Scope the DAG with selectors

mediapulse_base also contains a streaming domain that isn't yours today. Before you go anywhere near the lineage graph, confirm your actual scope on the command line - you've used --select with a single model before, now you need it for three domains at once:

dbt ls --select source:ads+ source:news+ source:podcasts+
  • What does source:podcasts+ select on its own, and what does the trailing + add that source:podcasts alone wouldn't?
  • Confirm nothing from streaming shows up anywhere in the output.
Hint: combining selectors

Selectors separated by spaces are unioned - source:ads+ source:news+ source:podcasts+ returns every node matched by any of the three, not just ones matched by all of them (that's what a comma, with no space, would do instead). + after a node means "and everything downstream of it" (staging → intermediate → marts); + before a node means upstream instead. See Set operators and Graph operators for the rest.

2. Walk the lineage

Open the lineage graph for mediapulse_base - it should match the node list you just got on the command line - and trace ads, news, and podcasts from source through staging to marts.

  • One staging model joins two sources at once, plus a seed. Find it, and note what each input contributes.
  • The seed it uses is the only seed in this project. Find it and work out why it's there instead of a static list in the model itself.
  • One raw source gets read twice, independently, on two separate paths through the DAG instead of once via its own staging model. Find it - this is exactly what Catalog's Source Fanout rule (see step 3) is warning you about.
Hint: streaming is not your problem today

You'll also see mediapulse_analytics depending on the streaming staging layer, plus one podcasts and one news staging model. That's real, but it's Group 2/3 territory once dbt Mesh is in play - note it and move on.

3. Read the Recommendations tab

Open Recommendations for mediapulse_base and filter to the Modeling category.

  • Confirm the Multiple Sources Joined flag matches the staging model you found in step 2.
  • Confirm the Source Fanout (or duplicate-source) flag matches the source you found being read twice.

Now filter to Testing. You should find a handful of Missing Primary Key Test flags - roughly half sit in podcasts/news. List which models they're on.

What Recommendations actually checks

Catalog's Recommendations are powered by dbt_project_evaluator and grouped into Modeling, Testing, Documentation, and Governance rules. See Project recommendations for the full rule list - you'll use more of these categories in later topics.

4. Find the dead ends

Still in the lineage graph: models/intermediate/ has more than one model with nothing downstream of it - built, tested, documented, and never ref()'d by a mart.

For each one you find, write one sentence on what it looks like it was meant to feed, and whether that's a wiring gap (an easy fix) or a sign the model doesn't have a home yet. You're not fixing these now - some get wired up for real later today and tomorrow.

5. Check source freshness

Look at the config: block on each source in news and podcasts - none currently define freshness or loaded_at_field. (ads doesn't even declare a source() - its staging models read the raw tables directly, so there's nothing to check there yet.)

In one sentence: why does that matter more for a source you don't control the load schedule of than for a mart you build yourself?

Hint: current freshness syntax

As of dbt 1.10, freshness and loaded_at_field are configured under config: at the source or table level, not as top-level source keys:

sources:
  - name: news
    config:
      freshness:
        warn_after: {count: 24, period: hour}
        error_after: {count: 48, period: hour}
      loaded_at_field: _etl_loaded_at
See Add sources to your DAG for the full config.

Deliverable: your selector command and what it confirmed about streaming (1), the multi-source-join model and the fanned-out source (2), your Missing Primary Key Test list (3), your dead-end intermediate models with a one-line plan each (4), and your freshness answer (5).

Done?

You've read a real project's lineage and test coverage through tooling instead of guessing from file names, and you've got a running list of gaps you'll close for real over the next few topics.