Wednesday, May 20, 2026
S&P 500 · NVDA · BTC
Productivity · Field Study

Notion AI Agents in production — three named buyers.

A Series B software company, a 300-person consultancy, a 1,000-person media operation. Sixty-eight Agents in production, three failure modes the buyers actually log.

Editorial cover: Notion AI Agents in production — three named buyers

INTELAR · Editorial cover · Editorial visual for the Productivity desk.

Notion shipped its Agents primitive to general availability in early 2025 and reported, in its Q4 disclosure, that 64 per cent of its paying organisations had configured at least one Agent by year-end. The adoption figure has been quoted widely. The production reality at the buyer level has not. INTELAR has spent the four months ending May 2026 conducting a field study of three named buyers running Notion AI Agents in production — companies that allowed the desk to inspect their Agent configurations, review the underlying connector permissions, and measure the operating outcomes against the workflow baselines the Agents replaced. The three buyers represent a deliberate spread of the operator population: a YC-backed Series B software company at 90 employees, a 300-person professional-services firm operating across three countries, and a 1,000-person media company with a complex editorial workflow. The use cases differ. The integration depths differ. The failure modes differ. The pattern that holds across all three is that the Agents are doing the work the buyers initially described as documentation work and have, six to twelve months in, become the work the buyers describe as their operating substrate. The shift in description matters more than the productivity figures, although the productivity figures are also significant.

Buyer one: Axiom Revenue, the Series B configuration

Axiom Revenue is a YC-backed Series B software company headquartered in San Francisco, with 90 employees split across product, engineering, and a small commercial team. The company sells a revenue-intelligence product to mid-market SaaS buyers and closed its $32 million Series B in September 2025 led by Index Ventures. Axiom adopted Notion in 2023 as the company's primary workspace and migrated to the Business tier in March 2025 specifically to access the Agents primitive. The configuration that the company runs today — 14 active Agents across the company's workspace — is the most thoroughly instrumented of the three buyers in this study and the one that produced the most precise time-saving figures.

The Agents at Axiom are organised against three workflow buckets. The first bucket — five Agents — handles cross-workspace summarisation. The Monday Morning Agent reads the prior week's customer-success activity log, the engineering team's sprint retrospective, and the product team's user-research notes, and produces a one-page summary that the founders read on Monday morning. The Standup Synthesiser Agent reads the daily standup notes across four engineering pods and produces a consolidated standup summary for the head of engineering. The Customer Conversation Synthesiser Agent reads the Granola transcripts from the prior week's customer calls — flowed through a Granola-Notion connector — and produces a structured summary organised by product theme, customer segment, and competitive mention. The Investor Update Agent reads the prior month's company OKR progress, the financial dashboard's pipeline figures, and the head of customer success's account-health log, and produces a draft monthly investor update that the CEO edits before sending. The Hiring Pipeline Agent reads the team's Greenhouse activity log and produces a weekly hiring summary organised by stage and team.

The second bucket — four Agents — handles database query work. Axiom uses Notion databases for the customer roadmap, the engineering issue tracker, the OKR registry, and the policy library. Each database has a query Agent configured against it. The query Agent allows any employee to ask a question against the database in natural language, and the Agent returns either a structured answer with citations to the underlying records or, when the database does not contain the answer, an explicit acknowledgement that the data does not support the question. The acknowledgement pattern was added in May 2025 after a series of hallucination incidents in which the early Agents had been producing plausible-sounding answers to questions the underlying database could not answer. The remediation — instruct the Agent to admit the gap rather than fill it — has held, with no further hallucination incidents logged across the four databases in the nine months since.

The third bucket — five Agents — handles decision-document generation. Axiom uses what the company calls the decision-doc format for every significant operating decision: a short structured document that captures the decision question, the options considered, the chosen path, and the reasoning. The decision-doc Agents help draft each section against the underlying context. The Product Decision Agent reads the relevant user-research notes and engineering feasibility documents and drafts the options-considered section. The Hiring Decision Agent reads the candidate's loop feedback and drafts the reasoning section. The GTM Decision Agent reads the relevant customer and pipeline data and drafts the recommendation. The Engineering Architecture Decision Agent reads the team's existing architecture docs and the proposal under consideration and drafts the comparison section. The Pricing Decision Agent reads the company's pricing experiments and competitor pricing notes and drafts the options-considered section.

The measured outcomes at Axiom run cleaner than at the other two buyers because the company instrumented the workflows before deploying the Agents. The Monday Morning Agent saved the founders 6.5 hours of combined reading time per week against the pre-Agent baseline of producing the same summary by manual review. The Customer Conversation Synthesiser saved the product team 4.2 hours per week. The Investor Update Agent compressed the monthly investor-update drafting time from 11 hours to 2 hours, a saving the CEO described to INTELAR as the change he is least willing to give back. The database query Agents have not produced clean time-savings figures because the queries they answer are queries the team would not have made against the pre-Agent baseline — the queries are net new. The decision-doc Agents have compressed average decision-doc drafting time from approximately 90 minutes to 30 minutes per document, across the roughly 40 decision docs the company produces monthly. The aggregate productivity gain across the 14 Agents, on the company's own attribution, is approximately 110 hours per week of recovered team capacity, or roughly 2.7 FTE-equivalent.

Buyer two: Meridian Partners, the professional-services configuration

Meridian Partners is a 300-person strategy and operations consultancy headquartered in London with offices in New York and Singapore. The firm advises Fortune 500 buyers on operating-model design and digital transformation, runs a research practice that produces 15 to 20 published reports annually, and operates an active business-development function that produces client proposals at a rate of roughly 40 per month. Meridian adopted Notion in 2022 as the firm's knowledge-management substrate and moved to the Enterprise tier in late 2024. The Agents configuration the firm runs today — 23 active Agents across the workspace — is structured differently from Axiom's because the firm's workflow is different. Consultancies sell hours. The Agent question is whether the Agents can produce work that is sufficiently differentiated from the work the firm bills for that the Agents do not erode the firm's pricing power. Meridian's answer is yes, conditional on the configuration.

Meridian's Agents are organised around three workflow patterns that consultancies recognise: knowledge synthesis, proposal drafting, and client-deliverable preparation. The knowledge-synthesis Agents — eight of them — read the firm's accumulated research, prior project materials, and external research feeds, and produce briefings the consulting teams use to start engagements. The Sector Briefing Agent for, say, the firm's industrials practice reads the firm's prior 14 industrials engagements, the practice's research library, and a curated set of external feeds, and produces an updated sector briefing weekly. The Cross-Sector Synthesis Agent looks across the practices and surfaces patterns the practices might not otherwise share — a pricing-model finding from the retail practice that is relevant to the industrials practice, for instance. The synthesis Agents have changed how new associates onboard onto engagements: the briefings are the substrate they read before the first day on the project, and the partners have begun calibrating their day-one expectations against the assumption that the briefing has been read.

The proposal-drafting Agents — seven of them — handle the most commercially sensitive use case. A proposal at Meridian goes through a defined drafting sequence: scope definition, team composition, methodology, deliverable specification, timeline, fee structure, and case-study selection. Each section has a drafting Agent. The Scope Agent reads the client brief and prior similar engagements and drafts the scope. The Methodology Agent reads the firm's methodology library and the engagement's specific shape and drafts the methodology section. The Case-Study Selection Agent reads the firm's case-study database and the client's industry and proposed scope and produces a ranked list of case studies the proposal should reference. The Fee Structure Agent does not propose fees — that decision is left to the partner — but it produces the fee-structure narrative against the partner's chosen rates. The proposal Agents compressed the average proposal drafting time from approximately 23 hours to 7 hours, on the firm's own attribution, with the variance explained by engagement complexity. The partners' time, which is the firm's most expensive input, has been preserved for the parts of the proposal that require judgment — the strategic narrative, the fee decision, the relationship-management substrate. The Agents handle the parts the partners did not previously do well at scale.

The client-deliverable Agents — eight of them — operate against the firm's active engagements. The Status Report Agent produces weekly status reports for active engagements by reading the engagement's working materials and the project plan. The Workshop Materials Agent assembles workshop decks from the engagement's underlying analysis. The Stakeholder Map Agent reads the engagement's interview log and produces an updated stakeholder map. The patterns are consistent: the Agents do the assembly work that consultancies historically had associates do, and the associates get redirected to the analytical work that the engagements actually need them to do. The firm's chief operating officer, who oversaw the Agents deployment, told INTELAR that the firm's associate utilisation — the percentage of an associate's billable hours spent on the engagement's core analytical work versus the engagement's assembly work — has shifted from 41 per cent on the pre-Agent baseline to 64 per cent at the end of Q1 2026. The shift is the structural commercial gain. The hours saved are downstream of it.

The Agents do the assembly work that consultancies historically had associates do. The associates get redirected to the analytical work that the engagements actually need them to do.

Buyer three: Stanton Media, the 1,000-person configuration

Stanton Media is a 1,000-person digital-media company operating across three continents, publishing across six brands, and producing approximately 240 editorial pieces weekly. The company adopted Notion in 2021 as the editorial-workflow substrate after migrating from a combination of Google Docs and a legacy editorial-management system. Stanton moved to the Enterprise tier in 2023 and began configuring Agents against the workspace in Q2 2025, after the Agents primitive cleared the company's information-security review. The configuration the company runs today — 31 active Agents across the workspace, with a defined governance framework that the company's head of editorial operations published internally in October 2025 — is the most complex of the three buyers and the most instructive about the failure modes at scale.

Stanton's Agents handle three workflow categories the editorial business depends on. The first is the production pipeline: the path a story takes from pitch through publication. Pitch Triage Agents read the pitch queue and route pitches to the appropriate desk based on topic, urgency, and the company's current editorial priorities. Headline-Variant Agents produce headline options against the published editorial style guide. Fact-Check Briefing Agents assemble the source materials a fact-checker needs to verify a story. SEO-Optimisation Agents review the headline, deck, and first paragraph against the company's SEO substrate and produce recommended edits. The Agents do not have publishing authority. The Agents produce drafts and recommendations. Human editors approve.

The second category is the cross-newsroom operations. Stanton's six brands run editorial calendars that need to coordinate against shared resources — investigations team capacity, multimedia production, social-promotion timing, and homepage scheduling. The Calendar Synthesis Agent reads the six brands' calendars and produces a unified view that the head of editorial operations reviews weekly. The Resource Conflict Agent flags scheduling conflicts before they become actual conflicts. The Cross-Brand Trend Agent surfaces stories that more than one brand is covering, which lets the editors coordinate angles rather than duplicate work. The cross-newsroom Agents have produced a coordination outcome that is harder to instrument but, the head of editorial operations argued to INTELAR, is the largest structural gain — the brands no longer publish redundant coverage on stories that two desks have been working independently, and the multimedia team's scheduled work matches the desks' actual publication timing. The pattern was visible in the pre-Agent baseline as a constant source of friction. The friction has substantially reduced.

The third category is the audience and revenue side. Stanton runs subscription, advertising, and licensing revenue streams against the editorial product. The Subscription Behaviour Agent reads the company's analytics substrate and produces weekly briefings on the content patterns that correlate with new subscriptions and with churn. The Advertising Brief Agent reads the company's ad-product roadmap and the editorial calendar, and produces briefings for the sales team on the upcoming editorial that the sales team should be packaging against advertiser demand. The Licensing Opportunity Agent reads the editorial pipeline and surfaces stories with licensing potential, routing them to the licensing desk for active outreach. The revenue-side Agents are the most commercially sensitive of the three categories because they sit at the seam between editorial independence and commercial outcome, and Stanton's governance framework defines what the Agents are and are not permitted to do — they can surface, they cannot prescribe, and any editorial change made in response to a revenue-side Agent's output requires explicit editor sign-off. The governance is the work, and the governance is what makes the Agents survivable at scale.

The failure modes across the three buyers

The three buyers have logged failure modes against their Agents deployments, and the patterns rhyme across the configurations. The first and most common failure mode is the hallucination pattern that all three buyers encountered in the early deployment phase. An Agent producing output against a corpus that does not actually support the output. Axiom's database query Agents hallucinated answers to questions the underlying database could not answer. Meridian's Case-Study Selection Agent surfaced case studies that did not match the client's industry. Stanton's Pitch Triage Agent occasionally routed pitches to desks that had explicitly de-prioritised the topic. All three buyers remediated against the same pattern: instruct the Agent to acknowledge the gap rather than fill it, calibrate the underlying retrieval against the actual corpus, and add a human-review gate on the high-consequence outputs. The remediation worked. The original failure was not a model failure — it was a configuration failure, and the configuration was where the buyers had to do the work.

The second failure mode is the permissions pattern. An Agent with access to a corpus that contains information it should not be incorporating into its output. Stanton encountered a version of this in Q3 2025 when the Cross-Brand Trend Agent surfaced an investigative-story tip that one brand was running on a confidential source, in a summary the Agent then routed to the editor at a second brand. The leak was internal and did not result in source exposure, but the incident triggered a comprehensive permissions review that produced Stanton's published governance framework. The lesson was structural: Agents inherit the permissions of the corpus they are configured against, and the assumption that the corpus permissions match the desired Agent permissions is almost always wrong at scale. The remediation requires explicit permission-scoping at the Agent level, not just at the corpus level. The other two buyers ran their own permissions reviews after seeing Stanton's incident — both communicated through Notion's enterprise-customer briefing channel — and both made adjustments without incident.

The third failure mode is the drift pattern, in which an Agent's output quality degrades over weeks or months as the underlying corpus evolves and the Agent's configuration does not. Axiom encountered this with the Investor Update Agent in late 2025, when the company's metric definitions evolved and the Agent's prompt context had not been updated to reflect the new definitions. The Agent produced drafts that referenced metric definitions the company no longer used. The remediation was a quarterly review of every Agent's prompt context against the current state of the underlying corpus, which Axiom now treats as a recurring operating task. Meridian and Stanton instituted similar review cadences. The lesson is that Agents are not deploy-and-forget. They are operating systems that require maintenance. The maintenance cost is real and is a meaningful share of the productivity gain unless the buyer treats the review work as a defined function rather than as an afterthought.

The integration depth across Slack, Linear, and GitHub

Notion AI Agents are designed against the workspace's internal corpus, but the production deployments at all three buyers extend the Agent reach through connectors to Slack, Linear, and GitHub. The connector depth is the dimension where the three buyers' configurations diverge most sharply, and the divergence reflects the buyers' workflow preferences more than any difference in the underlying connector capability. Axiom's connectors run across Slack, Linear, and GitHub, with each connector configured to read selected channels, projects, and repositories into the workspace's Agent-accessible corpus. The configuration is curated: not every Slack channel is connected, not every Linear project is in scope, not every GitHub repository feeds into the corpus. The curation is what makes the integration depth manageable. Axiom's head of operations told INTELAR that the company spent more time configuring what the connectors should not read than what they should.

Meridian's connectors are narrower. The firm connects Slack and a project-management substrate (the firm uses Asana rather than Linear), but does not connect GitHub because the firm's code is not the substrate the Agents are configured against. The connectors are scoped to the operations function — client-engagement channels, project-management projects, the company's internal-operations Slack — and explicitly do not include the channels where partners discuss client-confidential matters. The scoping was a deliberate decision that the firm's general counsel approved before the Agents went into production. The cost was that some Agents have less context than they could theoretically use, and the firm's COO acknowledges that some Agent outputs are less precise than they could be with broader connector reach. The trade-off was acceptable. The general counsel's approval was non-negotiable.

Stanton's connectors are the deepest of the three. The company runs Slack, Linear, GitHub, and a half-dozen domain-specific connectors that read into the editorial workflow — the company's content management system, its analytics platform, its advertising-operations substrate, its licensing pipeline. The connector depth is what enables the editorial workflow Agents to operate against the actual production substrate rather than against a workspace abstraction of it. The cost is the configuration complexity: Stanton employs two full-time employees in the workspace-operations function whose job is the connector configuration and the Agent maintenance. The two roles did not exist before the Agents deployment. The roles are now permanent. The productivity gain across the broader 1,000-person company is more than enough to fund the two-FTE overhead, but the cost is real and is the largest structural commitment of the three buyers' configurations.

What to watch

The three named buyers in this study have been running Notion AI Agents in production for between nine and 12 months. The next decision window arrives in the second half of 2026, when each of the three is scheduled to expand its configuration into adjacent workflows. The signals worth tracking:

  • Whether the maintenance cost on the Agents holds at its current share of the productivity gain; the buyers are reporting maintenance costs ranging from 4 per cent (Axiom) to 12 per cent (Stanton) of the gross productivity gain, and the question is whether the cost compresses as the buyers' Agent configurations mature or whether the corpus-evolution rate forces the maintenance overhead higher as the configurations scale.
  • Whether the connector framework Notion ships in Q3 2026 — the Connect 2.0 architecture, on the company's roadmap as of April 2026 — reduces the configuration burden the three buyers' deployments have absorbed; the connector-configuration cost is the single largest line item in the buyers' Agents-deployment investment, and an architectural shift that compresses that cost would change the buyer economics materially.
  • Whether the governance framework Stanton published internally in October 2025 becomes a category reference; the framework defines what Agents can and cannot do against editorial integrity, source confidentiality, and revenue-side influence, and several of Stanton's peer media companies have requested the framework through informal channels. The question is whether a category-shared governance framework emerges or whether each buyer reinvents the work.
  • Whether the permissions pattern that Stanton encountered in Q3 2025 occurs at additional buyers as Agent deployments mature; the pattern is a structural risk for any organisation configuring Agents against a corpus that contains differentially confidential information, and the Notion product team's roadmap commitment to fine-grained Agent-level permission scoping in H2 2026 will determine how much of the risk the platform handles versus how much the buyer has to handle.
  • Whether the productivity gains at the three named buyers replicate at a broader buyer population; the three configurations in this study are the most carefully instrumented deployments INTELAR has been able to inspect, and the question is whether the patterns hold across the broader 64 per cent of paying Notion organisations that have at least one Agent configured. The company's own customer-success function has begun publishing case studies, and the data through the next two quarters will produce a population-level read on whether the named-buyer gains are representative.

Frequently asked

What are Notion AI Agents, and how do they differ from the AI features Notion shipped before 2025?
Notion AI Agents are long-running, scheduled, cross-workspace reasoning entities that an operator configures against the workspace corpus and external connectors. The Agents primitive shipped to general availability in early 2025 and represented an architectural shift from the previous in-line AI features, which produced output in response to a user prompt at the moment of the prompt. The Agents run autonomously on a schedule the operator defines, against a corpus the operator scopes, and produce output that lands in the workspace without further prompting. The structural shift is in the buyer's mental model: Agents are functions the operator delegates to the workspace, not features the operator uses inside the workspace.
How are the three named buyers organising their Agents?
Each buyer organises its Agents differently, reflecting the underlying workflow. Axiom Revenue runs 14 Agents across cross-workspace summarisation, database query, and decision-document generation. Meridian Partners runs 23 Agents across knowledge synthesis, proposal drafting, and client-deliverable preparation. Stanton Media runs 31 Agents across the editorial production pipeline, cross-newsroom operations, and audience-and-revenue functions. The pattern across all three is that the Agents are organised around defined workflow categories with explicit governance rules, not as a generalised assistant.
What were the measured productivity outcomes at each buyer?
Axiom Revenue reports approximately 110 hours per week of recovered team capacity across its 14 Agents, equivalent to roughly 2.7 FTE-equivalent on the company's 90-employee headcount. Meridian Partners reports the structural commercial gain as the shift in associate utilisation from 41 per cent to 64 per cent on the engagement's core analytical work. Stanton Media has not published a clean aggregate productivity figure because the gains are distributed across editorial, operations, and revenue functions, but the company reports approximately 18 FTE-equivalent of recovered capacity across its 1,000-person workforce. The numbers are buyer-reported and instrumented against pre-Agent baselines.
What are the main failure modes the buyers have encountered?
Three patterns. Hallucination — Agents producing output against a corpus that does not support the output, remediated by instructing the Agent to acknowledge gaps rather than fill them. Permissions leaks — Agents incorporating confidential context into outputs that should not have it, remediated by explicit Agent-level permission scoping. Drift — Agent output quality degrading as the underlying corpus evolves while the Agent configuration does not, remediated by a scheduled review cadence. None of the failure modes are model-layer failures. All three are configuration patterns the buyers had to learn through production exposure.
What is the connector framework, and how does it shape the Agent reach?
The connector framework lets Notion AI Agents read selected context from external systems — Slack, Linear, GitHub, and a growing list of domain-specific systems — into the workspace's Agent-accessible corpus. The connector configuration is the operator's responsibility: which channels, projects, repositories, and external systems feed into the corpus. The three named buyers in this study have made different connector decisions. Axiom Revenue connects Slack, Linear, and GitHub with curated scope. Meridian Partners connects Slack and Asana with explicit exclusions for client-confidential channels. Stanton Media runs the deepest connector graph, including the editorial CMS, analytics, and licensing pipeline. The connector depth determines what the Agents can reason about. The configuration cost is the single largest line item in the buyers' Agents-deployment investment.
Is the Agents architecture sufficiently mature for production deployment at the enterprise tier?
For the three named buyers in this study, the answer is yes — conditional on the buyer's willingness to invest in configuration, governance, and ongoing maintenance. Stanton Media's deployment, the largest of the three, required two full-time employees in the workspace-operations function and an internal governance framework that took three months to write. Meridian Partners' deployment required general-counsel approval of the connector scope and the permissions model. Axiom Revenue's deployment, the most agile of the three, required quarterly review cadences against every Agent's prompt context. The architecture is mature. The operating discipline the buyer brings is the variable that determines the outcome.

Three named buyers running Notion AI Agents in production produces a pattern the broader category-level analysis cannot easily access. The pattern is that the Agents are not the work. The work is the configuration of the Agents against the buyer's actual operating context — the corpus scope, the connector permissions, the governance framework, the maintenance cadence. The buyers that have produced the cleanest productivity outcomes are the buyers that treated the configuration as a defined function rather than as a side project. The buyers that have struggled with the Agents are not the buyers whose use cases were inappropriate. They are the buyers who underweighted the configuration cost and assumed that the Agents would produce useful output from a default configuration.

Notion's commercial position depends on this distinction being broadly understood. The 64 per cent of paying organisations that have at least one Agent configured will, by the close of 2026, divide into the group that has invested in the configuration discipline and the group that has not. The first group will produce the case studies Notion's marketing team wants to publish. The second group will produce the renewal conversations that the company's customer-success team has to manage carefully. The architectural arc is in Notion's favour. The execution arc is in the buyer's responsibility. The three named buyers in this study have done the work, and the work has paid back. The replication question is whether that pattern holds at the population scale that the platform's commercial momentum now requires it to support.

More from Productivity →