Wednesday, May 20, 2026
S&P 500 · NVDA · BTC
Productivity · Field Notes

The weekly-review ritual reimagined with AI editors.

Eighty-four operators surveyed across six months. The four-part format the cohort converged on, the 0.7-point decision-quality differential, the failure modes the practice will need to survive.

Editorial cover: The weekly-review ritual reimagined with AI editors

INTELAR · Editorial cover · Editorial visual for the Productivity desk.

The weekly review is the operating practice the productivity literature claims everyone should do and almost no one does well. INTELAR has spent six months running a structured survey of an opt-in cohort of 84 chiefs of staff and heads of operations who have, since late 2024, been publishing about or quietly practising a version of the weekly review that incorporates AI editors as a working substrate. The cohort represents a deliberate spread: 31 chiefs of staff at venture-backed companies between Series B and pre-IPO, 22 heads of operations at the same stage, 18 founder-operators running pre-Series B companies who hold the chief-of-staff function themselves, and 13 senior partners at professional-services firms. The survey produces a clean read on a practice the trade press has been describing imprecisely. The format the cohort is converging on is more rigorous than the earlier weekly-review traditions — the David Allen GTD review, the founder's Friday close — and the convergence is happening because the AI editors have created a structural advantage the previous formats could not match. The cohort calls the practice many things. The function is consistent. The function is what this article will document.

The format the cohort has converged on

The weekly review that emerged from the INTELAR-surveyed cohort runs against a four-part structure that the 84 respondents have, with minor variation, all adopted. The four parts are calendar reflection, decision log, decision-readiness scoring, and a role-playing pre-mortem with an AI editor against the upcoming week. Each part runs for between 15 and 25 minutes. The whole review takes between 60 and 90 minutes. The cohort schedules the review at the same time each week — the modal Friday afternoon is the most common slot, with Sunday evening the second most common and Monday morning the third — and the practice is sufficiently established that 79 per cent of the cohort report not having missed a weekly review in the previous 12 weeks. The discipline is unusual. The format produced the discipline.

Calendar reflection is the first part. The operator opens the prior week's calendar and walks through it meeting by meeting. The AI editor — typically configured as a Claude Projects instance with persistent memory across the operator's weekly reviews, or as a Granola-integrated reflection workspace — receives the meeting list and a brief operator-written note on each meeting's outcome. The editor produces a structured summary that groups the meetings by category, surfaces patterns, and asks specific questions about the meetings the operator characterised as unproductive. The questions are not generic. They are produced against the operator's prior weekly reviews and the patterns the editor has identified across them. An operator who has been complaining about poorly run product reviews for three consecutive weeks gets a different set of questions than one whose product reviews have been productive. The persistent context is the structural feature the previous weekly-review formats could not produce.

The decision log is the second part. The operator writes down the significant decisions made during the prior week — typically between three and 11 decisions, with the cohort averaging seven — and the rationale for each. The AI editor reviews the decision log against the decisions made in prior weeks, flags decisions that appear to contradict earlier commitments, and asks the operator to either reconcile the contradiction or to document the reason for the change. The editor also flags decisions that the operator made without writing down the criteria that produced them, and prompts the operator to add the criteria post-hoc. The discipline this produces — the requirement to articulate the reasoning behind a decision rather than just the outcome — is what the cohort cites most often as the most valuable element of the review. The operator who can articulate the criteria can be challenged on them. The operator who cannot articulate the criteria cannot defend the decision when it comes up in a future review.

Decision-readiness scoring, the third part of the practice

Decision-readiness scoring is the part of the practice that the cohort has documented most variably and that produces the most consistent reported value. The mechanic is straightforward. The operator identifies the decisions the upcoming week will require them to make — typically five to nine decisions on the chief-of-staff and head-of-ops profile, somewhat fewer on the founder-operator profile, somewhat more on the senior-partner profile — and assigns each decision a readiness score from one to five. A score of five means the operator has the information they need, has resolved the trade-offs, and is ready to make the decision. A score of one means the operator does not yet know what question they are actually answering. The intermediate scores represent specific gaps the operator needs to close before the decision is ready: missing information, unresolved stakeholder alignment, an unconverted policy question, or the operator's own ambivalence about the criteria.

The AI editor's role in the decision-readiness scoring is to challenge the operator's self-assessment. The editor receives the decision list and the operator's scores, and produces a counter-assessment based on the operator's prior decision-readiness scoring against decisions of similar type. An operator who has historically over-scored their readiness on hiring decisions — making the decision and then revisiting it within two weeks — gets challenged on the hiring decisions in the current week. An operator who has historically under-scored their readiness on commercial decisions, then made the decisions confidently and accurately when the time came, gets less challenge on the commercial decisions. The pattern recognition is the value the editor adds. The operator's own assessment is reliable in some categories and unreliable in others, and the editor's persistent memory across the operator's prior reviews surfaces the unreliability the operator cannot see in themselves.

The decision-readiness scoring produces the measurable outcome that the cohort tracks most consistently. INTELAR's survey asked respondents to score, in retrospect, the quality of the decisions they made in the week following each weekly review against a five-point scale. The cohort that did not run weekly reviews — the control group of 47 chiefs of staff and heads of operations the survey also captured for comparison — scored their decisions at a 3.1 average across the survey window. The cohort that ran weekly reviews with the four-part structure including decision-readiness scoring scored their decisions at a 3.8 average across the same window. The 0.7-point differential is meaningful, although the methodology is not blinded and the self-reporting bias is real. The 3.8 average for the practising cohort is, however, internally consistent across the four buyer profiles, and the operators who scored their decisions highest in the practising cohort were the operators with the most consistent attendance at the weekly review. The correlation does not prove causation. The correlation is strong enough to inform the cohort's continued practice.

The operator who can articulate the criteria can be challenged on them. The operator who cannot articulate the criteria cannot defend the decision when it comes up in a future review.

The pre-mortem substrate, the fourth part

The fourth part of the weekly review is the most distinctively AI-editor-enabled and the part that the cohort has been most variable in adopting. The mechanic is the upcoming-week pre-mortem: the operator describes the upcoming week's most consequential meeting or decision, and the AI editor role-plays as one of the stakeholders the operator will encounter. The editor's prompt context includes the prior interactions with that stakeholder — meeting summaries from Granola, prior email threads if the operator chooses to include them, the operator's prior characterisations of the stakeholder's positions and motivations. The editor produces a simulated interaction in which the stakeholder pushes back on the operator's planned approach, raises objections the operator may not have anticipated, and proposes counter-arguments the operator has to respond to in the simulation.

The role-play does not produce a script. It produces a stress test. The operator who walks into the actual meeting after the pre-mortem has, in the cohort's reported experience, prepared responses for objections they would not have anticipated, has thought through positions they would not have rehearsed, and has identified the specific arguments the stakeholder is most likely to make. The value is not the editor's accuracy in simulating the stakeholder — the simulation is imperfect, and the operators acknowledge it — but the value is in the operator's pre-meeting preparation work the simulation forces. The operator who has done the work arrives in the meeting having already imagined the difficult turns. The operator who has not done the work is improvising. The improvisation cost compounds across the week.

The cohort's adoption of the pre-mortem substrate has been the part of the practice most affected by the operator's comfort with the AI editor format. The 18 founder-operators in the survey have been the heaviest adopters of the pre-mortem, citing the absence of a senior peer with whom to rehearse difficult conversations as the structural gap the editor fills. The 22 heads of operations have been the lightest adopters, with several reporting that they prefer to rehearse difficult conversations with a colleague rather than with an AI editor. The 31 chiefs of staff fall in the middle, with the adoption correlating most strongly with the operator's prior practice of preparing intensively for high-stakes meetings. The senior partners in professional services have been the most idiosyncratic adopters, with several using the pre-mortem to simulate client interactions and to rehearse the more difficult conversations a partner has to have with junior team members about performance. The pre-mortem is the most personal of the four parts. The cohort's variability reflects the personal nature.

The tools the cohort uses and how they configure them

The tools the surveyed cohort uses are concentrated in a small number of products that the operator population has converged on. Granola is the meeting-memory substrate that 78 per cent of the cohort reports using as the source for the calendar reflection part of the review. The operator who has been running Granola through the prior week opens the weekly review with a structured set of meeting summaries already produced, which removes the friction of reconstructing the week from raw memory. The Granola summaries flow into the operator's review workspace either through the company's Notion connector, through Slack, or through a direct Claude Projects integration that several cohort members have configured against their Claude API access. The integration depth is part of the structural feature: the AI editor does not need the operator to reconstruct the week, because the substrate has already captured it.

Reflect, the personal-knowledge-management tool, is the second most-cited substrate, used by 41 per cent of the cohort. Reflect's structural strength for the weekly review is the daily-notes pattern the tool encourages — the operator's daily notes for the prior week, including the spontaneous reflections the operator captured during meetings or while travelling, become an additional input to the review. Mem, the AI-first personal knowledge tool, is used by 23 per cent of the cohort against a similar structural advantage. The cohort members who use Reflect or Mem report that the value is the captured context the AI editor can read against the rest of the review, not the tool's own AI features in isolation. The substrate matters more than the model.

Claude Projects — the persistent-context Claude product — is the AI editor substrate that 64 per cent of the cohort reports using. The cohort members configure a Projects instance specifically for the weekly review, populate it with the operator's prior reviews as persistent context, and prompt the editor against the four-part structure. The persistence is the feature. The editor does not start fresh each week. The editor reads the prior 12 weeks of reviews, identifies the patterns, and challenges the operator against the patterns. ChatGPT's Projects equivalent and Notion AI Agents are the second and third most-cited AI editor substrates, each at lower adoption rates, with the cohort members citing Claude's longer-context performance and the cleaner separation between context and prompt as the operational reasons for the choice. The model layer is converging on Claude for this use case. The choice is functional rather than ideological.

The time investment and the payoff economics

The 60-to-90-minute weekly investment the practice requires is non-trivial for an operator population that already runs against time scarcity. The cohort's reported sustained adoption — 79 per cent attendance across 12 weeks — suggests the payoff exceeds the cost, but the payoff is structurally diffuse. The cohort cites three categories of return. The first is the decision-quality improvement that the 0.7-point differential in self-scored decision quality captures. The second is the reduced cost of stakeholder management — the cohort reports that the role-play pre-mortem produces meetings that conclude faster, with fewer rounds of follow-up, and with cleaner alignment than the meetings the cohort entered without the preparation. The third is the operator's own clarity on the week's priorities, which the cohort describes as the most subjective of the returns and the most consistent in the qualitative responses.

The decision-quality improvement is the most defensible return because it is measurable, although the measurement is imperfect. The cohort's self-scored decision quality improvement of 0.7 points on a five-point scale represents, the cohort estimates in the qualitative responses, a meaningful reduction in re-decision frequency. The operator who scores their decisions higher revisits them less often. The reduced revisitation is the time saving that compounds across the operator's week. INTELAR's survey did not capture the actual hours saved from reduced re-decision, but the cohort's estimates ranged between two and seven hours per week of recovered capacity from the cleaner decision-making. The 90-minute weekly review investment pays back, on the cohort's own attribution, between 1.3x and 4.6x against the recovered decision-making capacity alone.

The stakeholder-management return is harder to instrument but the qualitative cohort responses are consistent. Operators report that the pre-mortem-prepared meetings conclude with clearer decisions, fewer parking-lot items, and cleaner action-item assignment. The downstream effect — meetings the operator does not have to schedule because the prior meeting concluded cleanly — is the cost the cohort cites most often as the surprise. The pre-mortem produces a different meeting. The different meeting produces fewer follow-up meetings. The fewer follow-up meetings produce time on the operator's calendar that the operator can use for the work the review prepares them to do. The cycle is reinforcing. The reinforcement is what produces the 79 per cent attendance the cohort reports.

Failure modes and the operating discipline

The cohort has logged failure modes against the practice that the productivity literature should document. The first and most common failure mode is the editor-drift pattern. An AI editor that has been running against the operator's prior reviews can begin overweighting the patterns the editor has previously identified, to the point that the editor's questions become predictable and the operator's responses become rehearsed. The cohort that has run the practice longest — the founder-operators, several of whom have been running this format since late 2024 — reports the highest incidence of editor drift. The remediation is a quarterly reset of the editor's context, in which the operator either resets the Projects instance or actively prompts the editor to challenge the established patterns. The remediation is uncomfortable. The discipline requires it.

The second failure mode is the over-reliance pattern, in which the operator begins outsourcing the reflection work to the editor and going through the motions of the review without doing the underlying thinking. The cohort describes this as the format collapsing into ritual without reflection. The remediation is the operator's own discipline — the recognition that the editor is a substrate, not a substitute, and that the value of the review depends on the operator's willingness to do the difficult work the editor surfaces. The cohort members who have caught themselves in the over-reliance pattern report restoring the practice's value by deliberately disabling parts of the editor's automation for several weeks, returning to manual review against the format, and re-enabling the editor only after the manual practice has restored the thinking discipline. The pattern is recoverable. The discipline is the work.

The third failure mode is the corpus-context contamination pattern, which the cohort encounters when the editor's context begins drawing on information the operator did not intend to include in the review. An operator who has connected their Slack to the editor's context can find the editor surfacing items from team channels that the operator had not been thinking about in the review. The surfaced items are not always relevant, and the surface-noise cost can make the review less effective. The cohort's remediation is explicit scope management — defining what the editor's context should and should not include for the review specifically, separate from the editor's broader access to the operator's information substrate. The scope management is the configuration work the cohort cited as the most underestimated cost of the practice. The configuration time is real. The configuration time pays back. The cohort that has done the configuration runs the practice cleanly. The cohort that has not is producing reviews that the editor's context contaminates.

What to watch

The weekly-review-with-AI-editors practice has stabilised in the surveyed cohort but has not yet diffused widely into the broader operator population. The signals worth tracking in the next 12 months:

  • Whether the published-thought-leadership cohort produces a category-standard format the broader operator population can adopt; several of the surveyed operators have begun publishing about the practice on their own substacks and in operator-tier communities, and the question is whether the format converges into a published methodology or whether each operator continues to adapt the structure to their own context.
  • Whether the tooling layer produces a purpose-built product for the practice; Granola, Reflect, Mem, and Claude Projects are the substrates the cohort is using, but no product has been built specifically as a weekly-review-with-AI-editors product. The category demand is plausibly large enough to support a focused product, and the venture pipeline contains at least one early-stage company INTELAR is tracking that is positioning against the practice specifically.
  • Whether the decision-quality improvement the cohort reports holds in a more rigorous measurement design; the self-scored 0.7-point differential is meaningful but the methodology is not blinded, and the question is whether a more rigorous study — possibly run by an academic researcher partnering with the operator community — produces a comparable result or whether the cohort's improvement is partly a selection effect.
  • Whether the over-reliance failure mode produces visible quality regressions in cohorts that adopt the practice without the underlying discipline; the practice is more demanding than the format makes it look, and the question is whether the broader population that adopts it does so with the reflection-quality the cohort has maintained or whether the broader adoption produces a population of operators going through the motions without the underlying work.
  • Whether the role-play pre-mortem substrate becomes a category in its own right; several of the cohort members described the pre-mortem as the part of the practice they would most miss if they stopped, and the question is whether a separately purpose-built pre-mortem product could capture the value the cohort has been getting from the Claude Projects configuration. The product category is plausible. The execution is open.

Frequently asked

What is the four-part weekly review the cohort has converged on?
Calendar reflection, decision log, decision-readiness scoring, and a role-play pre-mortem with an AI editor against the upcoming week. Each part runs for between 15 and 25 minutes. The whole review takes between 60 and 90 minutes. The cohort schedules the review at the same time each week — modal Friday afternoon, with Sunday evening and Monday morning as secondary slots — and reports 79 per cent attendance across 12 weeks. The format is more rigorous than the earlier weekly-review traditions and is structured specifically around the persistent-context capabilities AI editors enable.
Who are the operators practising this format?
INTELAR's surveyed cohort comprises 84 operators across four profiles. Thirty-one chiefs of staff at venture-backed companies between Series B and pre-IPO. Twenty-two heads of operations at the same stage. Eighteen founder-operators running pre-Series B companies who hold the chief-of-staff function themselves. Thirteen senior partners at professional-services firms. The cohort is opt-in and self-selected, which limits the generalisability of the survey results, but the breadth of profiles produces a defensible read on a practice that has been emerging across multiple operator subpopulations simultaneously.
What tools is the cohort using to run the practice?
Granola is the meeting-memory substrate that 78 per cent of the cohort uses as the source for the calendar reflection. Reflect (41 per cent) and Mem (23 per cent) are the personal-knowledge-management substrates the cohort uses for daily context capture. Claude Projects (64 per cent) is the dominant AI editor substrate, with ChatGPT's Projects equivalent and Notion AI Agents as secondary options. The cohort's choice of Claude Projects is functional rather than ideological, citing the longer-context performance and the cleaner separation between context and prompt as the operational reasons. The model layer is converging on Claude for this use case.
What is the measured outcome the cohort reports?
A 0.7-point improvement in self-scored decision quality on a five-point scale, against a control group of 47 operators who do not run the practice. The methodology is not blinded and the self-reporting bias is real. The differential is, however, internally consistent across the four buyer profiles, and the operators who scored their decisions highest were the operators with the most consistent attendance at the weekly review. The cohort also reports between two and seven hours per week of recovered capacity from reduced re-decision frequency, which produces a payback ratio against the 90-minute weekly review investment of between 1.3x and 4.6x.
What are the failure modes the cohort has encountered?
Three patterns. Editor drift, in which the AI editor's pattern recognition becomes predictable and the operator's responses become rehearsed, remediated by quarterly context resets. Over-reliance, in which the operator outsources the reflection work to the editor and the format collapses into ritual without thinking, remediated by deliberately disabling parts of the editor's automation. Corpus-context contamination, in which the editor's context surfaces items the operator did not intend to include in the review, remediated by explicit scope management. The failure modes are recoverable. The discipline is the work.
Is the practice appropriate for every operator?
No. The practice is demanding enough that the operator's willingness to do the underlying reflection work determines the outcome. Operators who adopt the format without the reflection discipline produce reviews that collapse into ritual within four to eight weeks. The practice is most defensible for operators whose weeks contain multiple consequential decisions, whose stakeholder management is complex enough that the pre-mortem substrate produces value, and whose prior practice already included some version of structured reflection. The cohort that has sustained the practice is the cohort that brought the reflection discipline. The format amplifies the discipline. The format does not create it.

The weekly review with AI editors has produced a format the operator population is converging on, a measurable decision-quality improvement the surveyed cohort can defend, and a set of failure modes that the practice's broader adoption will need to work through. The cohort's experience suggests that the AI editor substrate produces a structural advantage the earlier weekly-review traditions could not match — the persistent context across reviews, the pattern recognition across an operator's prior assessments, the role-play pre-mortem that the operator could not previously rehearse with another human at the necessary frequency. The advantage is real. The discipline the practice requires is also real. The operators who have been able to maintain the discipline have produced the outcomes the survey captures. The operators who have not maintained the discipline have not.

The category implication, beyond the practice itself, is that the AI editor substrate has begun to do work the operator population previously did with a senior peer or did not do at all. The pattern is visible in other operator practices — the call prep that Granola does, the meeting summary that Notion AI Agents do, the code review that Claude Code does — but the weekly review is the most distinctively operator-cognitive of the practices, because the substrate is the operator's own thinking rather than an external corpus. The cohort's experience suggests that the operator's own thinking, properly externalised, becomes the substrate against which the AI editor produces the most useful work. The implication for the broader operator population is that the practice is replicable. The implication for the cohort is that the practice has compounded across the past 18 months and continues to compound. INTELAR will revisit the cohort in early 2027 to measure whether the compound has held.

More from Productivity →