> ## Site Index
> Fetch the site index at: https://www.credibledata.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Curate What Your Agents See

> An agent with more permission than its caller is a leak. With less, it's a toy. Model the access rule and it inherits exactly theirs.

Your company already decided who may see what. Somebody wrote that policy years ago, and wrote it carefully. Now it has to cover things that aren't people.

Take one line of that policy: *support sees their own customers' tickets and nobody else's*. It exists as a row filter in the warehouse and again in the BI tool's permission model, and now an agent shows up and someone asks you to write it a third time. Each copy is a translation of the policy on its own release schedule, and nothing proves the three still agree. Every migration drops one copy and adds another.

A data leader at a large software company told us this year that their biggest problem wasn't the warehouse, the BI tool or AI. It was one set of ACLs -- everyone extracts data into their own tools, and no two analyses of the same question agree. Their answer was one platform for everyone. Ours is closer to the opposite: one control, enforced in one place, and everything else free to stay decentralized.

The usual fix is row-level policies on the tables themselves. It works, but it doesn't stay simple. You end up with one rule per table instead of one per decision, and nobody reviews a policy they can't read in one sitting.

**Declare the control next to the thing it governs.** For analytical data, that's your semantic data model. It's how everyone reaches the data, and it's the place that says what the data means. Write the rule on the same object as the definition and the two can't drift apart the way a policy document and the code that implements it always do. [Model the Meaning First](/blog/posts/model-meaning-first) made this argument for the pipeline. [The Semantic Layer Is an Interface, Not a Compiler](/blog/posts/semantic-layer-is-an-interface) made it for the API -- one of the five jobs an API does is enforce, at one gate instead of N.

The next post in this pair covers where controls come from: how a policy becomes a numbered list of testable things, and what that lets you prove before an agent is ever deployed. This one is about the control itself. It shows how to build a model where two annotations on one source are the entire access rule, and why two is enough.

## An agent only respects what the system enforces

This stopped being a slow-moving problem because the newest consumer of your data doesn't behave like the old ones.

A person who joins the analytics team gets the tour. They learn that the finance extract isn't for sharing, that the churn table has been wrong since the migration, that nobody touches the HR schema. They soak up a decade of institutional caution from the people sitting around them, and most of it was never written down anywhere.

An agent gets none of that. It reads what's there and believes it, and it'll traverse every table it can reach at machine speed, whether it runs in a chat window, an IDE or a product you shipped to customers. Nobody can tell it which schema was supposed to be off limits, so the only limits it has are the ones the system enforces.

The governance you already have covers this, as long as it's actually enforced. That enforcement is now the only thing between a question and an answer -- and once it's real, you can stop rationing the agent.

## The data model, in three layers

The control needs somewhere to sit, so here's the shape of a Credible data model first. We'll read it top down, in the order you'd model it.

In Malloy, a `source` is a named thing you can query: a table, plus whatever joins and definitions are built on it. A data model is a set of sources that build on each other.

**Analytical domains.** The only sources anyone queries, and one per access decision. A source carries one gate, so it gives one answer, and support, the exec team and the app in your product each need their own. In practice that's one per category of question, split wherever two audiences need different answers to it. This is the layer a person, a dashboard or an agent names. It's the interface, and it's where the access rule will go.

`is ticket_core extend { ... }` derives this one from the layer below and adds to it, so everything the core already knows comes along without being restated. A `view` is a saved query shape on the source: ask for `backlog_by_priority` and you get open tickets broken out by priority, without writing the grouping again.

```malloy
// support_performance.malloy
import "core.malloy"

source: support_performance is ticket_core extend {
  view: backlog_by_priority is {
    group_by: priority
    aggregate: open_ticket_count
  }
}
```

**The core.** The join network, and every definition more than one domain depends on. Nobody queries it directly.

`join_one` attaches a related source and says there's at most one of it per row on this side, so no question ever writes that join again. `dimension` names an attribute and `measure` names an aggregate. `open_ticket_count` carries its own filter, and that filter is `is_open` from the base source instead of another copy of `status = 'open'` -- which is the point of the layers. These are the numbers two teams would otherwise each define slightly differently. Support's first response time counts the first internal note and the exec team's doesn't, and both dashboards are right according to their own file.

```malloy
// core.malloy
import "base.malloy"

source: ticket_core is tickets_raw extend {
  join_one: account is accounts_raw on account_id = account.account_id
  join_one: assignee is agents_raw on assigned_agent_id = assignee.agent_id

  dimension:
    tenant is account.tenant_slug
    segment is account.segment

  measure:
    ticket_count is count()
    open_ticket_count is count() { where: is_open }
    avg_first_response_hours is first_response_hours.avg()
}
```

**Base sources.** One per physical table: shape, plus whatever a single table can say about itself.

`conn.table(...)` points at a table in the warehouse, and `primary_key` declares the grain (one row per ticket), which is what lets Malloy keep a `join_one` honest and aggregate correctly across it. Dimensions and measures belong here too, as long as they need nothing but this table's own columns: unit conversions, flags, the cleanups you'd otherwise repeat in every query. The join is the dividing line. Anything that has to reach another table isn't a property of this one.

```malloy
// base.malloy
source: tickets_raw is conn.table('support.tickets') extend {
  primary_key: ticket_id

  dimension:
    first_response_hours is first_response_minutes / 60
    is_open is status = 'open'
}

source: accounts_raw is conn.table('crm.accounts') extend {
  primary_key: account_id

  dimension:
    tenant_slug is lower(tenant_name)
}

source: agents_raw is conn.table('support.agents') extend {
  primary_key: agent_id
}
```

If that shape looks familiar, it's [MVC](https://en.wikipedia.org/wiki/Model%E2%80%93view%E2%80%93controller) with the names changed: base sources are the model, the core is the controller, analytical domains are the views. The mapping holds where it matters, because authorization goes on the view and parameters flow down into what sits beneath it.

Three layers is more setup than one flat file. In exchange, callers only ever start from the top layer, so that's the only place anyone has to decide who may see what. The [Modeling Overview](/docs/how-to/modeling/ai-modeling#the-shape-of-a-data-model) walks through the same three layers.

## Curate what can be queried

Before there's any rule, decide what's visible at all. It's the same `public`, `private` and `internal` decision an object-oriented codebase makes about its members, and it goes where OO puts it -- on every class.

Prefix a definition with `public`, `internal` or `private`, or set the levels in an `include` block on the source. What changes by layer is what the decision is *for*.

At the base, visibility hides the table. A raw table is mostly things nobody should build on: `status_v2`, the ETL batch id, three nullable columns nobody has touched since the migration, and the raw inputs to the definitions you just wrote. Once `first_response_hours` exists, `first_response_minutes` is an implementation detail. Make the concepts public and the columns behind them private, and the layers above can use `first_response_hours` without ever seeing `first_response_minutes` or any other warehouse column name.

The core does the same thing one level up. Its join network needs scratch dimensions that exist only to make a measure come out right, and those are `internal` by definition. What the core makes `public` is the vocabulary the domains are allowed to build from.

At the domain, `public` means something different. It's the promise to callers, and it's the surface the gate will guard:

```malloy
// support_performance.malloy
import "core.malloy"

source: support_performance is ticket_core include {
  private: *
  internal: first_response_hours
  public: tenant, segment, priority, status, created_at,
          ticket_count, open_ticket_count, avg_first_response_hours
} extend {
  view: backlog_by_priority is {
    group_by: priority
    aggregate: open_ticket_count
  }
}
```

`private: *` closes the surface by default, so a column someone adds to the table next quarter stays hidden until somebody opts it in. `internal` keeps `first_response_hours` available to `avg_first_response_hours`, which is defined in terms of it, without making the per-ticket number something a caller can ask for. `public` lists the eight fields a caller can ask for, and that's the whole interface.

Below the top layer nothing can be queried directly, so visibility there has nothing to do with security. It decides what the next layer up may build on and keeps a lower layer's columns from leaking upward by accident. Only at the domain does it become the contract with callers.

Visibility is also static. It can't read a given, so a core can't show one set of fields to one domain and a different set to the next. It publishes the union and each domain narrows it -- which is why two audiences that need different columns end up as two sources.

Fields are one grain of visibility. The other is which sources anyone outside can see at all, and that's a separate decision made in one file. Each package has one `index.malloy`, and what it exports is the menu.

```malloy
// index.malloy
import "support_performance.malloy"

export { support_performance }
```

What it exports is everything a caller can name. The core and the base sources still compile, import and join. They just aren't somewhere a caller can start. The menu has nothing to do with identity: an agent over MCP, a person in a workspace, a dashboard, a data app and an API call from your own product all see the same one.

It's also enforced. A source off the menu answers 404 to every caller, and changing the menu takes a new published version, like any other change to the model. That's how you keep a table reachable through a join without letting anyone query it directly. A package with no `index.malloy` has no menu, and every source in it can be queried.

Between the two, you now have a short, named list of entry points, each with a known set of fields. That list is where the gate goes. [Curating Discovery](/docs/how-to/modeling/metadata-tags#curating-discovery) has the full visibility and menu rules.

## Two annotations, on the interface

Here's the rule, as two lines above the domain source from the top of the post:

```malloy
// support_performance.malloy
import "core.malloy"

given:
  #(secure)
  GROUPS :: string[]
  #(secure)
  TENANTS :: string[]

#(authorize) 'support' in $GROUPS
#(access_filter) tenant in $TENANTS
source: support_performance is ticket_core include {
  private: *
  internal: first_response_hours
  public: tenant, segment, priority, status, created_at,
          ticket_count, open_ticket_count, avg_first_response_hours
} extend {
  view: backlog_by_priority is {
    group_by: priority
    aggregate: open_ticket_count
  }
}
```

Curation decides what exists for anyone. Authorization decides who may use it. A source that's internal, or open to everyone, only needs the first:

| Mechanism | Decides | If refused |
| --- | --- | --- |
| **Curation** (same for every caller) | | |
| `public`, `internal`, `private` | which fields a source exposes | compile error |
| `index.malloy` | which sources can be queried | 404 |
| **Authorization** (depends on who's asking) | | |
| `#(authorize)` | whether the caller may query it | 403 |
| `#(access_filter)` | which rows the caller sees | 200, their rows |

### Where the rule lives

On the sources in `index.malloy` that need one, and nowhere else. Some exported sources are open to everyone who can reach the package and carry no rule. Others hold data only some people may see, and carry `#(authorize)`, `#(access_filter)` or both. Sources off the menu don't need a rule, because no caller can start a query from them. That's MVC's rule about authorization: it goes on the endpoints the world can see, not on every model class and helper beneath them.

This works because of one fact about gates: **a gate is checked on the source a query starts from, and it never follows a join.** If an open source joins a gated one, the gate doesn't fire, and the open source hands the gated data to anyone who can query it. So you protect the sources a caller can start from, not each sensitive table. Those are exactly the sources `index.malloy` exports, which is why curation came before the rule.

It also catches a common mistake. Somebody needs a number on a Thursday, writes a new source that joins the gated one, and publishes. In a flat model that source is live the moment it ships. Here it answers 404 until someone adds it to `index.malloy`, which takes a new published version, the same step as changing the rule itself.

One limit follows from the same fact: the rule covers queries that go through the model, and only those. A BI tool pointed straight at the warehouse, or an analyst with a SQL client, never touches the model, so none of this applies to them. Warehouse permissions still decide who can go around the model.

The same fact is why the top layer has one domain per access decision. Count them: a few hundred tables, and a dozen or so real decisions, every one of which your business already made and wrote down somewhere. Tie the rules to the tables and you write hundreds. Tie them to the decisions and you write a dozen.

*The gate is checked where a query enters and never follows a join, so the line is the only place policy can live. Widen a single entry point and one admission reaches everything joined beneath it; split the entry points and each admission reaches a slice.*

### What the rule says

There are two annotations because the line of policy this post opened with makes two decisions. `#(authorize)` is the first half, *support sees*: may this caller reach the source at all? `#(access_filter)` is the second, *their own customers' tickets and nobody else's*: which of its rows do they see? The filter applies to every query against the source, and no query can widen or remove it. Conditions that must all hold are joined with `and`, or written as repeated annotations. There's no `or`, so an either-or rule takes two sources. A source extended from a gated one inherits its gate unless it declares its own, which replaces it.

The policy goes into the model once, on the object that already says what a ticket is. There's no copy in a policy language, a permission model or a prompt to keep in sync.

### Who is asking

The givens are where identity comes in, and both of these are **secure givens**: Credible fills them server-side from the caller's verified identity and ignores anything the caller sends. `$GROUPS` is built in. It holds the caller's group names, or the group an API key acts as, and Credible always supplies it. `$TENANTS` is a custom attribute, resolved the same way from values assigned per user or per group. A custom secure given is only enforced if it's set-valued (`string[]`) and the rule reads it with a membership test; a scalar one isn't enforced. Every other kind of given comes from the caller. That's fine for parameterizing a query, and it can never be an access boundary.

The engine resolves identity once for whoever is asking: a person in a workspace, a dashboard, a data app, your product's API, an agent over MCP. Every one of them hits the same two lines carrying their own `$GROUPS` and `$TENANTS`. Adding a consumer or an agent adds no rule, so the next agent surface you connect is one more caller, and nobody has to rewrite the policy for it. [Access Control](/docs/how-to/modeling/fine-grained-acls#who-is-asking-secure-givens) covers assigning values to a secure given.

### What comes back

Three things can happen, and they look different on purpose. A source the menu doesn't reach is 404: for that caller, it doesn't exist. A source the gate refuses is 403, a refusal that's recorded like any other query. The [security](/docs/platform-admin/security) and [monitoring](/docs/platform-admin/monitoring) pages cover that record, and both are part of the [Enterprise plan](/pricing). A caller the gate admits gets 200 and their own rows.

The middle one matters, because it's tempting to return an empty result instead, and you shouldn't. For an aggregate, `where: false` doesn't return no rows. It returns a row anyway, with `count()` at 0 and `sum(salary)` as NULL. That's an answer about data the caller was refused, and it looks exactly like an ordinary empty result. A lock misconfigured against a whole tenant then reads as a data-quality problem until somebody files a ticket. The other failures go the same direction: a gate that can't be applied denies, and a malformed gate fails to publish instead of shipping with the protection missing.

## Try it on one domain

All it takes to start is one domain and an afternoon. Pick a category of question your team actually asks, one with a real rule attached to it. Write the base sources for its tables, sticking to each table's own concepts and nothing that needs a join. Build one core with the joins and the definitions everyone argues about. Put one domain on top, `include` only what it should expose, and list it in `index.malloy`. Then write the policy you already have, the line with the department name in it, as `#(authorize)` and `#(access_filter)` above the domain. Publish, and ask it a question as someone who shouldn't be able to see the answer. That's the whole loop, and the second domain takes an hour. The steps are in [Secure One Domain](/docs/how-to/modeling/fine-grained-acls#secure-one-domain), and the rest of that page is the reference for both annotations.

## One control, on the model

A gate evaluates against whoever is asking. When an agent asks on your behalf, it gets your rows. Let a colleague in another region use the same agent and it gets theirs, and nothing about the answer looks filtered. That's the difference between an agent you can give the whole company and one you can only give three people. Give it more permission than its user has and you've built a data leak with a chat interface. Give it less and it's a toy, and the people who need real answers go back to extracting CSVs. The one worth deploying has exactly the caller's permissions -- and you never had to decide what the agent may see. You decided what each person may see long before anyone mentioned agents.

Every consumer, human or agent, passes through the same two lines at the one gateway [Inside the AI Analytics Engine](/blog/posts/inside-the-ai-analytics-engine) describes, and every query is recorded there with who asked, from which surface, against which model version. What keeps an agent from reading what it shouldn't is the same rule that scopes the dashboard, checked before the question is answered and recorded after.

The model is the only thing in your stack that lasts as long as the decisions do. Warehouses get migrated, BI tools get replaced, and the agent surface you integrate this quarter isn't the one you'll run in three years. The decisions about what your data means and who may see it outlive every one of those tools, and they shouldn't have to be made again each time a tool gets swapped out. [Charting is Commoditized. Meaning Is the Product.](/blog/posts/meaning-is-the-product) argued that the tools are disposable and the meaning isn't. An access rule is part of the meaning.

The data leader who told us their problem was one set of ACLs wanted one platform for everyone. The answer turned out to be smaller than that: one control, on the model, and everything else free to stay where it is. Somebody wrote your policy carefully, years ago, and has watched it get retyped ever since. Give it somewhere to live that outlives your vendors.
