Support
Log In

Modeling Overview

Build data models with AI agents — in the app or in your IDE

A data model is Malloy code, in a versioned package, that defines what your data means: sources over your tables, the joins between them, and the dimensions, measures, and views that encode your business definitions. The model is the input. The engine derives the materialized tables, access enforcement, and agent context from it.

You build models with an AI agent — in the app with Build & Publish, or in your IDE with the developer tools. Both produce the same packages, and this section applies to both.

New to data modeling? See The Data Model for the concepts, or the Malloy Language Documentation for language details.

What Goes in the Model

Put everything that defines your data in the model, so there is no second copy to keep in sync:

  • Entities and relationships, with their cardinality
  • Metrics, each defined once
  • Business rules and edge cases — trial churn doesn't count, intercompany transfers are excluded, the fiscal year starts in February
  • Descriptions of what each field means

You express these with a handful of Malloy constructs:

ConstructWhat it does
SourceA table from a connection, extended with definitions: source: orders is conn.table('sales.orders') extend { ... }
JoinA relationship between sources, declared once and available to every query
DimensionAn attribute to group and filter by
MeasureAn aggregate, such as revenue or order count
ViewA saved query built from dimensions and measures
Visibilitypublic, internal, and private decide which fields a source exposes; index.malloy decides which sources callers can query
AnnotationA tag that tells the engine how to treat a source or field — #(doc), #(index), #(authorize), #(access_filter), #@ persist

For example:

#@ persist name="order_metrics"
source: order_metrics is conn.table('orders') extend {
  join_one: customers is conn.table('customers') with customer_id
  join_many: items is conn.table('order_items') on order_id = items.order_id

  dimension:
    #(doc) Drives onboarding campaign eligibility.
    customer_segment is customers.signup_date ?
      pick 'new' when >= now - 30 days
      else 'established'

  measure:
    #(doc) Revenue at order placement. Recognized revenue lives in finance.recognized_revenue.
    booked_revenue is sum(order_total)
    order_count is count()
}
  • Cardinality is declared. join_one and join_many state that each order has one customer and many line items. The compiler enforces it, so joined totals never double-count — see Why Malloy.
  • Rules are defined once. customer_segment is written here, not repeated in every query.
  • Names carry distinctions. booked_revenue, not revenue, with a #(doc) pointing to recognized revenue, so an agent can't confuse the two.
  • Storage is an annotation. #@ persist is the only mention of storage. The engine builds the table and keeps it fresh.

Agents receive these definitions as context, so an agent and a dashboard asking the same question get the same answer.

The Shape of a Data Model

A small model fits in one source. A larger one organizes its sources in three layers:

LayerWhat it holdsWho queries it
Analytical domainsOne source per access decision, derived from the core, with viewsEvery caller — people, dashboards, data apps, and agents
The coreThe joins, and every definition more than one domain usesNobody directly
Base sourcesOne source per table: its grain, plus definitions that need only its columnsNobody directly

Callers only query the top layer, so that is where access rules go. The structure mirrors MVC: base sources are the model, the core is the controller, and domains are the views.

The examples below model a support team's tickets, from the bottom layer up.

Base Sources

One source per table. primary_key declares the grain — one row per ticket — which lets Malloy validate joins and aggregate correctly across them. Add dimensions and measures that need only this table's columns, such as unit conversions and flags. Anything that reaches another table belongs in the core.

// base.malloy
source: tickets_raw is conn.table('support.tickets') extend {
  primary_key: ticket_id

  dimension:
    first_response_hours is first_response_minutes / 60
    is_open is status = 'open'
}

source: accounts_raw is conn.table('crm.accounts') extend {
  primary_key: account_id

  dimension:
    tenant_slug is lower(tenant_name)
}

source: agents_raw is conn.table('support.agents') extend {
  primary_key: agent_id
}

The Core

The joins, and the definitions that must mean the same thing in every domain. join_one attaches a related source with at most one match per row. open_ticket_count reuses is_open from the base source instead of repeating status = 'open'.

// core.malloy
import "base.malloy"

source: ticket_core is tickets_raw extend {
  join_one: account is accounts_raw on account_id = account.account_id
  join_one: assignee is agents_raw on assigned_agent_id = assignee.agent_id

  dimension:
    tenant is account.tenant_slug
    segment is account.segment

  measure:
    ticket_count is count()
    open_ticket_count is count() { where: is_open }
    avg_first_response_hours is first_response_hours.avg()
}

Analytical Domains

The sources callers query. Each carries at most one access rule, so create one domain per access decision — typically one per category of question, split wherever two audiences need different answers. is ticket_core extend inherits everything the core defines, and a view saves a query shape so callers don't rewrite it.

// support_performance.malloy
import "core.malloy"

source: support_performance is ticket_core extend {
  view: backlog_by_priority is {
    group_by: priority
    aggregate: open_ticket_count
  }
}

To control which fields each layer exposes, and which domains callers can reach, see Curating Discovery.

Build a Model with the Agent

Before you start, you need:

The workflow is the same in the app and in your IDE:

  1. Describe what you want to model. Be specific about your data and goals:

    • "Build a model of my ecommerce data so I can analyze sales by product and brand"
    • "Create a data model for customer analytics including lifetime value"
    • "Model the orders table with customer and product relationships"
  2. The agent explores and proposes. Using its MCP tools, the agent explores your indexed tables and proposes which to include, how they join, and which dimensions and measures to define. Credible's open-source agent skills guide it to follow Malloy best practices on every surface.

  3. Confirm and iterate. Approve or adjust each proposal in plain language. The agent writes the model, with #(doc) on every field and #(index) where it helps discovery.

Migrating from a BI tool or semantic layer? If your business logic lives in Looker, Power BI, Tableau, Cube, dbt, or a warehouse semantic view, give the agent those definitions instead of starting from scratch. It rebuilds them as Malloy, which is faster and more faithful than describing the model from memory. See Migrations.

Validate as You Build

Check results against real data before you publish:

  • In the app — the agent runs queries against the draft package and shows results in the chat. Open the draft at any time to browse the generated files.
  • In your IDE — use the buttons above each source: Schema shows the compiled structure, Explore opens the interactive Explorer, and Preview runs a quick data check.

Changing the Model

A change to a definition is made in one place, and everything derived from it — tables, docs, indexes, and access enforcement — moves with it in the next published version. Versions and rollback are covered in Publishing. To let domain experts propose changes for review, see Git-Backed Modeling.

Stages of Model Development

The pages in this section follow a model's stages in order:

  1. Model — sources, joins, dimensions, measures, and views (this page)
  2. Curate and document — decide which sources and fields are visible, then document and index them with Discovery
  3. Secure — decide who may query each source and which rows they see, with Access Control
  4. Optimize — materialize expensive sources with Performance & Cost
  5. Publish — version and serve the model to every consumer with Publishing

A first model can go straight from stage 1 to publishing, then add the rest as it evolves.

Next Steps

On this page