Support
Log In

Discovery

Curate, document, and index your model so people and agents find and understand your data

Discovery decides what people and agents can find in your model, and how well they understand it. It has two steps:

  1. Curate — decide which sources and fields are visible at all.
  2. Document — describe what each visible field means with #(doc), and make dimension values searchable with #(index).

When you publish, the engine compresses your #(doc) definitions and #(index) values into the concept index, which people and agents search by meaning.

The agent adds #(doc) and #(index) to every field it defines as part of the modeling workflow, and can document an existing model the same way. Use this page to review its work. For how indexes are built, kept fresh, and paid for, see Performance & Cost.

Curating Discovery

Curation works at two grains, both in the model:

GrainMechanismControls
Fieldspublic, internal, privateWhich fields each source exposes
SourcesThe package's index.malloyWhich sources callers can query

Set both on every layer of the three-layer model, not just the top one. The examples below use the support model from that page.

Field Visibility

Prefix a definition with an access modifier, or set levels in an include block before extend:

ModifierQueryableUsable in definitions
public✅✅
internal❌✅ — in this source and sources that extend it
private❌✅ — in this source only

See Access Modifiers in the Malloy documentation for the full rules.

What to expose at each layer:

LayerMake publicMake internal or private
Base sourcesThe table's concepts, such as first_response_hoursRaw columns and inputs to derived fields, such as first_response_minutes, ETL ids, and deprecated columns
The coreThe definitions domains build fromHelper fields that exist only to compute a measure
Analytical domainsThe fields callers may queryEverything else

At the domain, the public fields are the source's interface — what callers can query and what the access rule guards:

// support_performance.malloy
import "core.malloy"

source: support_performance is ticket_core include {
  private: *
  internal: first_response_hours
  public: tenant, segment, priority, status, created_at,
          ticket_count, open_ticket_count, avg_first_response_hours
} extend {
  view: backlog_by_priority is {
    group_by: priority
    aggregate: open_ticket_count
  }
}
  • private: * hides every field by default, so a column added to the table later stays hidden until it is listed.
  • internal: first_response_hours keeps the field available to avg_first_response_hours without making it queryable.
  • public: lists the eight fields callers can query. Querying any other field is a compile error.

An include block can also be a denylist — include { except: customer_email, credit_card_number } — but use the allowlist form for sensitive data, so new columns aren't exposed automatically.

Visibility is static: it can't vary by caller. To show a column to some callers only, use two sources.

The Menu: index.malloy

Each package has one index.malloy. The sources it exports are the only ones callers can query:

// index.malloy
import "support_performance.malloy"

export { support_performance }

A source that isn't exported still compiles and can be imported, joined, and extended, but it is:

  • Not listed — people and agents browsing the package don't see it
  • Not indexed — its #(doc) and #(index) values never enter the concept index or value search
  • 404 by name — a query that names it fails as if the source doesn't exist

The menu is the same for every caller, on every surface. Changing it requires publishing a new version.

A package without an index.malloy has no menu, so every source in it can be queried. Add one to any package with layers beneath its domains.

Curation controls discovery, not access: what exists for everyone, not who may query it. For per-caller rules, see Access Control. A caller refused by #(authorize) doesn't see that source in their agent's context, so you don't need the menu to hide it.

Documentation Tags

Put #(doc) on the line before a field to describe what it means:

source: orders is conn.table('sales.orders') extend {
  dimension:
    #(doc) Customer segment based on lifetime value: High (>$10k), Medium (>$1k), Low
    customer_segment is case
      when lifetime_value > 10000 then 'High'
      when lifetime_value > 1000 then 'Medium'
      else 'Low'
    end

  measure:
    #(doc) Revenue at order placement. Recognized revenue lives in finance.recognized_revenue.
    booked_revenue is sum(amount)
}

Write descriptions that add what the code can't say:

  • Include units, thresholds, business rules, caveats, what depends on the definition, and where related numbers live.
  • Don't restate the field name or logic the model already expresses.
  • Name the distinction, too. Agents take field names at face value, so call booked revenue booked_revenue, not revenue.

Value Indexing

Put #(index) on a dimension to index its distinct values, so agents can find the field by searching for a value, not just a field name or description:

source: products is conn.table('catalog.products') extend {
  dimension:
    #(index)
    #(doc) The product category (e.g., Electronics, Clothing, Home & Garden)
    category is product_category

    #(index)
    #(doc) The brand name
    brand is brand_name
}

For example, if product_category holds values like "Running Shoes" and "Athletic Apparel", a question about "sports gear" matches nothing by name. With #(index), the engine matches the question to those values and helps the agent filter on them.

IndexSkip
Names and titles — products, customers, programsNumeric fields — amounts, counts, IDs
Categorical values — statuses, types, categoriesTimestamps and dates
Lookup values — regions, departments, brands

Value search is off on access-controlled sources, because the index is shared across callers. See Value Search and Materialization.

Complete Example

A source with documentation and value indexing applied:

source: orders is conn.table('sales.orders') extend {
  primary_key: order_id

  join_one: customers is conn.table('sales.customers') on customer_id = customers.id
  join_one: products is conn.table('catalog.products') on product_id = products.id

  dimension:
    #(doc) Date the order was placed
    order_date is created_at::date

    #(doc) Order status: pending, processing, shipped, delivered, cancelled
    #(index)
    status is order_status

    #(doc) Product category from the catalog
    #(index)
    category is products.product_category

    #(doc) Customer's geographic region
    #(index)
    region is customers.region

  measure:
    #(doc) Total number of orders
    order_count is count()

    #(doc) Total revenue in USD
    total_revenue is sum(amount)

    #(doc) Average revenue per order in USD
    avg_order_value is total_revenue / order_count

    #(doc) Percentage of orders that were cancelled
    cancellation_rate is count() { where: status = 'cancelled' } / order_count * 100

  view:
    #(doc) Monthly revenue trend with order counts
    monthly_revenue is {
      group_by: order_date.month
      aggregate: total_revenue, order_count
    }
}

Next Steps

On this page