Discovery
Curate, document, and index your model so people and agents find and understand your data
Discovery decides what people and agents can find in your model, and how well they understand it. It has two steps:
- Curate — decide which sources and fields are visible at all.
- Document — describe what each visible field means with
#(doc), and make dimension values searchable with#(index).
When you publish, the engine compresses your #(doc) definitions and #(index) values into the concept index, which people and agents search by meaning.
The agent adds #(doc) and #(index) to every field it defines as part of the modeling workflow, and can document an existing model the same way. Use this page to review its work. For how indexes are built, kept fresh, and paid for, see Performance & Cost.
Curating Discovery
Curation works at two grains, both in the model:
| Grain | Mechanism | Controls |
|---|---|---|
| Fields | public, internal, private | Which fields each source exposes |
| Sources | The package's index.malloy | Which sources callers can query |
Set both on every layer of the three-layer model, not just the top one. The examples below use the support model from that page.
Field Visibility
Prefix a definition with an access modifier, or set levels in an include block before extend:
| Modifier | Queryable | Usable in definitions |
|---|---|---|
public | ✅ | ✅ |
internal | ❌ | ✅ — in this source and sources that extend it |
private | ❌ | ✅ — in this source only |
See Access Modifiers in the Malloy documentation for the full rules.
What to expose at each layer:
| Layer | Make public | Make internal or private |
|---|---|---|
| Base sources | The table's concepts, such as first_response_hours | Raw columns and inputs to derived fields, such as first_response_minutes, ETL ids, and deprecated columns |
| The core | The definitions domains build from | Helper fields that exist only to compute a measure |
| Analytical domains | The fields callers may query | Everything else |
At the domain, the public fields are the source's interface — what callers can query and what the access rule guards:
// support_performance.malloy
import "core.malloy"
source: support_performance is ticket_core include {
private: *
internal: first_response_hours
public: tenant, segment, priority, status, created_at,
ticket_count, open_ticket_count, avg_first_response_hours
} extend {
view: backlog_by_priority is {
group_by: priority
aggregate: open_ticket_count
}
}private: *hides every field by default, so a column added to the table later stays hidden until it is listed.internal: first_response_hourskeeps the field available toavg_first_response_hourswithout making it queryable.public:lists the eight fields callers can query. Querying any other field is a compile error.
An include block can also be a denylist — include { except: customer_email, credit_card_number } — but use the allowlist form for sensitive data, so new columns aren't exposed automatically.
Visibility is static: it can't vary by caller. To show a column to some callers only, use two sources.
The Menu: index.malloy
Each package has one index.malloy. The sources it exports are the only ones callers can query:
// index.malloy
import "support_performance.malloy"
export { support_performance }A source that isn't exported still compiles and can be imported, joined, and extended, but it is:
- Not listed — people and agents browsing the package don't see it
- Not indexed — its
#(doc)and#(index)values never enter the concept index or value search - 404 by name — a query that names it fails as if the source doesn't exist
The menu is the same for every caller, on every surface. Changing it requires publishing a new version.
A package without an index.malloy has no menu, so every source in it can be queried. Add one to any package with layers beneath its domains.
Curation controls discovery, not access: what exists for everyone, not who may query it. For per-caller rules, see Access Control. A caller refused by #(authorize) doesn't see that source in their agent's context, so you don't need the menu to hide it.
Documentation Tags
Put #(doc) on the line before a field to describe what it means:
source: orders is conn.table('sales.orders') extend {
dimension:
#(doc) Customer segment based on lifetime value: High (>$10k), Medium (>$1k), Low
customer_segment is case
when lifetime_value > 10000 then 'High'
when lifetime_value > 1000 then 'Medium'
else 'Low'
end
measure:
#(doc) Revenue at order placement. Recognized revenue lives in finance.recognized_revenue.
booked_revenue is sum(amount)
}Write descriptions that add what the code can't say:
- Include units, thresholds, business rules, caveats, what depends on the definition, and where related numbers live.
- Don't restate the field name or logic the model already expresses.
- Name the distinction, too. Agents take field names at face value, so call booked revenue
booked_revenue, notrevenue.
Value Indexing
Put #(index) on a dimension to index its distinct values, so agents can find the field by searching for a value, not just a field name or description:
source: products is conn.table('catalog.products') extend {
dimension:
#(index)
#(doc) The product category (e.g., Electronics, Clothing, Home & Garden)
category is product_category
#(index)
#(doc) The brand name
brand is brand_name
}For example, if product_category holds values like "Running Shoes" and "Athletic Apparel", a question about "sports gear" matches nothing by name. With #(index), the engine matches the question to those values and helps the agent filter on them.
| Index | Skip |
|---|---|
| Names and titles — products, customers, programs | Numeric fields — amounts, counts, IDs |
| Categorical values — statuses, types, categories | Timestamps and dates |
| Lookup values — regions, departments, brands |
Value search is off on access-controlled sources, because the index is shared across callers. See Value Search and Materialization.
Complete Example
A source with documentation and value indexing applied:
source: orders is conn.table('sales.orders') extend {
primary_key: order_id
join_one: customers is conn.table('sales.customers') on customer_id = customers.id
join_one: products is conn.table('catalog.products') on product_id = products.id
dimension:
#(doc) Date the order was placed
order_date is created_at::date
#(doc) Order status: pending, processing, shipped, delivered, cancelled
#(index)
status is order_status
#(doc) Product category from the catalog
#(index)
category is products.product_category
#(doc) Customer's geographic region
#(index)
region is customers.region
measure:
#(doc) Total number of orders
order_count is count()
#(doc) Total revenue in USD
total_revenue is sum(amount)
#(doc) Average revenue per order in USD
avg_order_value is total_revenue / order_count
#(doc) Percentage of orders that were cancelled
cancellation_rate is count() { where: status = 'cancelled' } / order_count * 100
view:
#(doc) Monthly revenue trend with order counts
monthly_revenue is {
group_by: order_date.month
aggregate: total_revenue, order_count
}
}