Orply.

Data Agents Need Reviewed Semantic Context, Not Just Table Access

OpenAIFriday, September 18, 20264 min read

OpenAI argues that a useful data agent needs the same context a new analyst would receive, not just access to a warehouse. In its ChatGPT Work demonstration, the company shows how to combine Databricks tables, Slack discussions and metric documentation into a reusable semantic layer, then have a data expert review its definitions and rules before sharing it as a company plugin. The aim is to make the organization’s interpretation of metrics available alongside the underlying records.

A data agent needs the context a new analyst would receive

Giving an agent access to data is not the same as giving it the context to analyze that data correctly. The speaker’s test is simple: consider what a new analyst joining the team would need in order to work responsibly, then give that material to the agent.

What would I give a new analyst on the team? Give that to the agent.

That framing makes the task broader than connecting a model to a warehouse. A new analyst needs to know which tables matter, how the team talks about the business, which metric definitions are authoritative, and what filters or denominator rules govern recurring analysis. Without that context, the same raw data can support incompatible interpretations.

In the example, that context becomes a data-context skill called the “Sites Semantic Layer.” The displayed document describes its area as Sites adoption, activation, creation-to-publication, and repeat usage. Its intended users are Sites analysts and data scientists working with the relevant demo tables. Rather than functioning as a single answer to a single question, the layer is a reference meant to shape how future analyses are performed.

The layer also captures rules that sit alongside the raw records. Its “Standard Filters and Dimensions” section lists rules, default logic, and the metrics to which they apply. “External human users,” for core adoption and lifecycle metrics, are defined by user_type = 'human' joined to an organization condition of org_type == 'external'. For creator-denominator and eligible-user metrics, weekly eligibility is defined as eligible_from <= week_start.

Those rules establish who belongs in a metric population and when a user counts as eligible. If an analyst—or an agent—uses a different definition of an external user, or applies eligibility at a different point in time, it can produce a different adoption or creator-rate result from the same underlying tables. The operational purpose of the semantic layer is to make the team’s intended interpretation available alongside the data, so that its defaults can be applied consistently rather than rediscovered in each request.

The agent must connect records, team knowledge, and metric definitions

The business requirement is to connect the systems that hold the evidence for an analysis, not merely the system that holds the records. In the demonstration, the speaker first ensures that ChatGPT Work is connected to the same tools they would use to do the job themselves. The Data plugin is then pointed to three kinds of material: Sites data tables in Databricks, an analytics channel in Slack, and a Google Drive document containing information about Sites metrics.

3
source types used to seed the data-context skill

Each source serves a distinct role in the resulting context. The Databricks tables provide the underlying records. The Slack analytics channel supplies working team context: the discussions in which questions, conventions, and potentially relevant information are held. The Google Drive document supplies metric information in a more formal reference. The prompt shown in the interface asks the Data plugin to use those supplied assets and any other relevant information it finds to create the context layer.

This is why the example does not treat a data connection as sufficient. A table may contain columns and events, but it does not by itself state which populations should be excluded from a core lifecycle metric, which users form a denominator, or how the team understands “activation” and “repeat usage.” The displayed semantic layer is the place where those analytical choices are assembled into an explicit working reference.

The mechanics of setup remain straightforward in the source: users can go to the Plugins tab, select the plus button to add a plugin, and log into an account where authentication is required. The speaker notes that business and enterprise users may need an administrator to enable particular plugins. But those permissions are a means to the larger requirement: the agent needs access to the tools and materials from which a human analyst would construct an informed interpretation.

The resulting “Sites Semantic Layer” brings the seeded materials together in one artifact. It is presented as a skill that can be reviewed and reused, rather than as a one-off query run against connected systems. That distinction matters because the recurring value is not only an answer to a current request; it is a shared set of context and analytical defaults available for subsequent work.

Generated context still needs an accountable owner

The generated layer is not presented as the final authority. The speaker says someone from the data team who knows, or is familiar with, the data should review and refine it. The agent can look across the connected sources and assemble a draft, but a data expert is expected to decide whether the resulting context accurately represents the team’s data and metric logic.

That review is especially consequential for the parts of the layer that convert business terms into operational rules. The displayed definitions for external human users and weekly eligibility are compact, but they govern the populations used in core adoption, lifecycle, creator-denominator, and eligible-user metrics. Packaging an unreviewed interpretation would risk spreading a rule that looks reusable without reflecting the team’s intended definition.

Once the data reviewer thinks the skill is in good shape, they can package it as a plugin and share it across the company. The reusable asset is therefore a reviewed data-context skill: one that combines access to the relevant tools with the team’s data, discussion, metric documentation, and analytical rules.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free