Support
Log In

Performance & Cost

Materialize sources, index dimensions, and pre-aggregate measures for fast, cheap, searchable models

Credible keeps managed, derived copies of your data to make queries faster and cheaper. You opt parts of your model in with an annotation, and the engine builds each copy, keeps it fresh, reuses it across versions, and serves from it.

Derived copyMakesAnnotate
Materialized tableA source fast and cheap to queryThe source, with #@ persist
Search indexA dimension's values searchableThe dimension, with #(index)
Pre-aggregationA measure fast and cheap at coarse grainsThe measure, with #@ preaggregate

A derived copy changes how fast a query returns, never what it returns. For what to index for AI retrieval, see Discovery; this page covers how indexes are built and served.

Deciding What to Persist

Materialize a source

Queried often or expensive to compute? Add #@ persist.

Index a dimension

Values people or agents search or filter by? Add #(index).

Pre-aggregate a measure

A hot measure queried at coarse grains? Add #@ preaggregate with its grain.

Leave it live

No annotation: queried live from the source database, not searchable.

For an expensive, frequently queried source, persist the source, index its dimensions, and pre-aggregate its hot measures. The bare annotation is usually enough: the defaults let the engine deduplicate copies across versions and schedule refreshes. Add a freshness window only for data with a real staleness requirement.

An indexed dimension may be partitioned by at most one required filter, and may not sit on a source that requires parameters. Both are publish-time errors.

Access-controlled sources have limits on materialization; see Value Search and Materialization.

Serving Behavior

A query serves from a derived copy when one covers it, and otherwise runs live against the source database — slower, never wrong.

  • A materialized table stores a source's data as a physical table and routes queries to it.
  • A search index makes a dimension's values findable by value search, filter suggestions, and agents.
  • A pre-aggregation stores a measure rolled up to a declared grain, and answers queries at that grain or any coarser one it can correctly re-aggregate to.

The copies build on each other automatically. On a materialized source, indexes and pre-aggregations are built from the materialized table rather than the warehouse, so:

  • Values are consistent — the index matches what queries return.
  • Refreshes are aligned — a table-backed index refreshes when its table does. An index on an unmaterialized source refreshes on publish, on demand, and on its freshness window if you set one.
  • Builds are cheaper — no second warehouse scan.

Credible sequences the builds, so an index or rollup is always built after the table it derives from.

Configuration

Annotations

Options are optional key="value" pairs on the annotation; omit them to accept the defaults.

#@ persist name="orders_fast" refresh="incremental" watermark="order_date" freshness.window="24h"
source: orders is conn.table('sales.orders') extend {
  #(index)
  dimension: status is order_status
}
OptionSets
nameWhere the materialized table lands (container-qualifiable)
refresh"full" (default) rebuilds the whole copy; "incremental" applies only new rows — see Incremental Refresh
watermark / merge_keyHow an incremental table finds and applies new rows
freshness.windowThe staleness objective the engine schedules against — see Freshness

Package Manifest

Reuse scope and refresh cadence are set once per package in publisher.json (see Publishing):

{
  "scope": "package",
  "materialization": { "freshness": { "window": "24h", "fallback": "live" } }
}
  • scope: "package" (default) — a derived copy is reused across versions that define the same thing. Lowest cost.
  • scope: "version" — each version keeps its own copies. Use it to own an exact rebuild schedule for a version.

Set either a freshness objective or a materialization.schedule, not both. A fixed schedule requires scope: "version".

Incremental Refresh

A full refresh recomputes the whole table. For a large, append-mostly source, refresh="incremental" reads only rows newer than the last build:

#@ persist refresh="incremental" watermark="order_date"
source: daily_revenue is orders -> {
  group_by: order_date
  aggregate: revenue is amount.sum()
}

The source body and queries against it are unchanged. Search indexes are always incremental.

KeyMeans
refresh="incremental"Advance the table with a delta instead of a full recompute. Requires watermark.
watermarkThe output dimension a refresh derives its range from, such as an event time or order date. Its values must be monotone — a row's value never decreases.
merge_keyOnly when a row's watermark moves (for example, updated_at). The row's stable identity — one or more comma-separated output dimensions — so a changed row replaces its old copy.

Publishing fails with a targeted error if a required key is missing, a named dimension doesn't resolve, or a key names a measure.

Choosing a Shape

SourceDeclareExample
Rollup or append-only factwatermark only. Each refresh replaces its range, including rows deleted upstream within it#@ persist refresh="incremental" watermark="ingested_at"
Mutable tablewatermark and merge_key#@ persist refresh="incremental" watermark="updated_at" merge_key="id"
No monotone dimension, such as a small lookup tableNothing — leave refresh unset#@ persist

Limits

Two gaps are reported at publish rather than handled automatically:

  • Late data — a row arriving with a watermark below the range already covered is never picked up.
  • Hard deletes — a deleted row never appears in a delta, so a merge_key source keeps it. Prefer soft deletes: keep the tombstone flag in the persisted output and filter on it downstream.

To repair either, correct the row upstream and advance its watermark, or force a full rebuild with Rerun on the package page (or forceFullRebuild on the runs API). Any model change triggers a full rebuild automatically.

Non-additive measures such as count_distinct and median are safe in incremental sources: affected rows are recomputed from full input. Windows that look forward along the watermark (lead(), whole-partition percentages) are rejected at publish; trailing windows are fine.

Pre-Aggregations

Mark a hot measure with its rollup grain, and Credible maintains a rolled-up copy and answers coarse queries from it. No query names the rollup; the engine routes automatically.

#@ preaggregate grain="order_time.day, category"
measure: total_revenue is amount.sum()
  • grain (required) — the dimensions the rollup stores. A query uses the rollup when everything it groups and filters by is covered, including coarser truncations of a stored time dimension (order_time.month over a day grain). Otherwise it runs live.
  • #@ -preaggregate — pins a measure to the base source even when a rollup covers it, for consumers that can't tolerate the rollup's freshness.

Routing never returns a wrong number:

MeasureServed from the rollup
Additive (sum, count, min, max) and avgAt any covered grain
Non-additive (count(distinct), median, percentiles)Only at exactly the declared grain, with a publish-time warning

Measures with the same grain share one rollup table. A rollup is built from the materialized table when the base is persisted, and from the warehouse when it isn't — often the right choice when a full copy of the base isn't worth keeping. A rollup is never fresher than its base; a stale rollup is skipped, not served.

Freshness

freshness.window is an objective, not a fixed refresh time: it sets how stale a derived copy may get, and the engine schedules refreshes to meet it — batching work, running off-peak, and skipping refreshes a recent publish or run already covered.

fallback sets what happens when a materialized table is older than its window: live runs the query against the warehouse instead of serving stale data.

Search indexes report their staleness on the version page and in search and retrieval responses, so agents can tell when suggestions come from an older snapshot.

Builds and Refreshes

WhenWhat happens
On publishEvery persisted source, index, and pre-aggregation for the new version is built
On demandRerun on the package page (or the runs API) rebuilds a version, a single source, or a single dimension — optionally with its upstream persisted sources
On scheduleCopies refresh to meet their freshness objectives

The version page lists Materialized sources and Indexed dimensions with their status and build history.

What It Costs

Derived copies meter as described on the pricing page:

  • Storage — tables, rollups, and indexes, billed per GB-month
  • Compute time — queries the engine answers from them, billed per second
  • Queries run directly on your warehouse — no compute charge from Credible

Storage Reclamation

Credible garbage-collects unused derived copies: a copy is kept only while an unarchived version references it. Archive versions you no longer use to reclaim their storage — auto-archive, on by default, does this on a retention window you control.

Next Steps

On this page