Modeling Overview
Build data models with AI agents — in the app or in your IDE
A data model is Malloy code, in a versioned package, that defines what your data means: sources over your tables, the joins between them, and the dimensions, measures, and views that encode your business definitions. The model is the input. The engine derives the materialized tables, access enforcement, and agent context from it.
You build models with an AI agent — in the app with Build & Publish, or in your IDE with the developer tools. Both produce the same packages, and this section applies to both.
New to data modeling? See The Data Model for the concepts, or the Malloy Language Documentation for language details.
What Goes in the Model
Put everything that defines your data in the model, so there is no second copy to keep in sync:
- Entities and relationships, with their cardinality
- Metrics, each defined once
- Business rules and edge cases — trial churn doesn't count, intercompany transfers are excluded, the fiscal year starts in February
- Descriptions of what each field means
You express these with a handful of Malloy constructs:
| Construct | What it does |
|---|---|
| Source | A table from a connection, extended with definitions: source: orders is conn.table('sales.orders') extend { ... } |
| Join | A relationship between sources, declared once and available to every query |
| Dimension | An attribute to group and filter by |
| Measure | An aggregate, such as revenue or order count |
| View | A saved query built from dimensions and measures |
| Visibility | public, internal, and private decide which fields a source exposes; index.malloy decides which sources callers can query |
| Annotation | A tag that tells the engine how to treat a source or field — #(doc), #(index), #(authorize), #(access_filter), #@ persist |
For example:
#@ persist name="order_metrics"
source: order_metrics is conn.table('orders') extend {
join_one: customers is conn.table('customers') with customer_id
join_many: items is conn.table('order_items') on order_id = items.order_id
dimension:
#(doc) Drives onboarding campaign eligibility.
customer_segment is customers.signup_date ?
pick 'new' when >= now - 30 days
else 'established'
measure:
#(doc) Revenue at order placement. Recognized revenue lives in finance.recognized_revenue.
booked_revenue is sum(order_total)
order_count is count()
}- Cardinality is declared.
join_oneandjoin_manystate that each order has one customer and many line items. The compiler enforces it, so joined totals never double-count — see Why Malloy. - Rules are defined once.
customer_segmentis written here, not repeated in every query. - Names carry distinctions.
booked_revenue, notrevenue, with a#(doc)pointing to recognized revenue, so an agent can't confuse the two. - Storage is an annotation.
#@ persistis the only mention of storage. The engine builds the table and keeps it fresh.
Agents receive these definitions as context, so an agent and a dashboard asking the same question get the same answer.
The Shape of a Data Model
A small model fits in one source. A larger one organizes its sources in three layers:
| Layer | What it holds | Who queries it |
|---|---|---|
| Analytical domains | One source per access decision, derived from the core, with views | Every caller — people, dashboards, data apps, and agents |
| The core | The joins, and every definition more than one domain uses | Nobody directly |
| Base sources | One source per table: its grain, plus definitions that need only its columns | Nobody directly |
Callers only query the top layer, so that is where access rules go. The structure mirrors MVC: base sources are the model, the core is the controller, and domains are the views.
The examples below model a support team's tickets, from the bottom layer up.
Base Sources
One source per table. primary_key declares the grain — one row per ticket — which lets Malloy validate joins and aggregate correctly across them. Add dimensions and measures that need only this table's columns, such as unit conversions and flags. Anything that reaches another table belongs in the core.
// base.malloy
source: tickets_raw is conn.table('support.tickets') extend {
primary_key: ticket_id
dimension:
first_response_hours is first_response_minutes / 60
is_open is status = 'open'
}
source: accounts_raw is conn.table('crm.accounts') extend {
primary_key: account_id
dimension:
tenant_slug is lower(tenant_name)
}
source: agents_raw is conn.table('support.agents') extend {
primary_key: agent_id
}The Core
The joins, and the definitions that must mean the same thing in every domain. join_one attaches a related source with at most one match per row. open_ticket_count reuses is_open from the base source instead of repeating status = 'open'.
// core.malloy
import "base.malloy"
source: ticket_core is tickets_raw extend {
join_one: account is accounts_raw on account_id = account.account_id
join_one: assignee is agents_raw on assigned_agent_id = assignee.agent_id
dimension:
tenant is account.tenant_slug
segment is account.segment
measure:
ticket_count is count()
open_ticket_count is count() { where: is_open }
avg_first_response_hours is first_response_hours.avg()
}Analytical Domains
The sources callers query. Each carries at most one access rule, so create one domain per access decision — typically one per category of question, split wherever two audiences need different answers. is ticket_core extend inherits everything the core defines, and a view saves a query shape so callers don't rewrite it.
// support_performance.malloy
import "core.malloy"
source: support_performance is ticket_core extend {
view: backlog_by_priority is {
group_by: priority
aggregate: open_ticket_count
}
}To control which fields each layer exposes, and which domains callers can reach, see Curating Discovery.
Build a Model with the Agent
Before you start, you need:
- A place to build — a workspace in the Credible App, or an IDE with the Credible Extension
- An indexed database connection — see Connect a Database. The agent sees a connection's tables only after it is indexed, which you set up in the connection's scope. Indexing can take a few minutes; if tables are missing, wait and check that they're within the indexing limits.
The workflow is the same in the app and in your IDE:
-
Describe what you want to model. Be specific about your data and goals:
- "Build a model of my ecommerce data so I can analyze sales by product and brand"
- "Create a data model for customer analytics including lifetime value"
- "Model the orders table with customer and product relationships"
-
The agent explores and proposes. Using its MCP tools, the agent explores your indexed tables and proposes which to include, how they join, and which dimensions and measures to define. Credible's open-source agent skills guide it to follow Malloy best practices on every surface.
-
Confirm and iterate. Approve or adjust each proposal in plain language. The agent writes the model, with
#(doc)on every field and#(index)where it helps discovery.
Migrating from a BI tool or semantic layer? If your business logic lives in Looker, Power BI, Tableau, Cube, dbt, or a warehouse semantic view, give the agent those definitions instead of starting from scratch. It rebuilds them as Malloy, which is faster and more faithful than describing the model from memory. See Migrations.
Validate as You Build
Check results against real data before you publish:
- In the app — the agent runs queries against the draft package and shows results in the chat. Open the draft at any time to browse the generated files.
- In your IDE — use the buttons above each source: Schema shows the compiled structure, Explore opens the interactive Explorer, and Preview runs a quick data check.
Changing the Model
A change to a definition is made in one place, and everything derived from it — tables, docs, indexes, and access enforcement — moves with it in the next published version. Versions and rollback are covered in Publishing. To let domain experts propose changes for review, see Git-Backed Modeling.
Stages of Model Development
The pages in this section follow a model's stages in order:
- Model — sources, joins, dimensions, measures, and views (this page)
- Curate and document — decide which sources and fields are visible, then document and index them with Discovery
- Secure — decide who may query each source and which rows they see, with Access Control
- Optimize — materialize expensive sources with Performance & Cost
- Publish — version and serve the model to every consumer with Publishing
A first model can go straight from stage 1 to publishing, then add the rest as it evolves.