Helix Vertical, a Series C compliance-automation company headquartered in Austin and operating across four time zones, runs a 40-person revenue-operations function under a single VP. Twelve months ago that team was 47 people, and the gap between the two numbers is not an attrition story. It is a workflow story. Beginning in March 2025, the VP — Lena Sokolowski, who joined from Stripe's revenue operations practice in Q4 2024 — instrumented her team's daily work against a defined set of agent workflows, measured the minutes recovered per workflow, and used the recovered capacity to hold headcount flat as the company grew bookings 41 per cent year over year. The headcount that did not get added is the productivity gain. The 12-month measurement window closed on 2 May 2026, and the data Sokolowski shared with INTELAR — under a named-source agreement that covers the workflow architecture and the minutes-saved figures but excludes the underlying revenue figures — produces the clearest field study of agent-assisted revenue operations the desk has seen this cycle. The team kept its full audit log. The audit log is the article.
The Helix Vertical context, in numbers
Helix Vertical sells a compliance-monitoring product to mid-market financial services and insurance buyers. The company closed its Series C — $58 million, led by Coatue with participation from Greylock — in November 2024, at a $410 million post-money valuation. By the time Sokolowski's instrumentation programme began in March 2025, the company carried 311 enterprise logos, an annualised recurring revenue of $34 million, and an average sales cycle of 87 days from first conversation to closed paper. The revenue-operations team supports the full motion: lead enrichment, account research, opportunity hygiene, renewal forecasting, and churn-risk scoring across the installed base. Five workflows. Forty people. The ratio matters.
The team's pre-March 2025 baseline was unremarkable for a company of Helix's stage. Lead enrichment ran on ZoomInfo, with manual verification on every account above $50,000 in opening ARR potential. Account research ran on internal templates produced in Notion and copy-pasted from a rotating cast of secondary sources. Opportunity hygiene ran on weekly forecast calls with a Salesforce-native pipeline review. Renewal forecasting ran on a quarterly cohort exercise using Looker-fed spreadsheets. Churn-risk scoring did not, at that point, run in any structured way at all. The function was operating well below its potential because the people doing the work were operating at the limit of their time, not their capability. Sokolowski's instrumentation hypothesis was that agent workflows could lift the time ceiling without raising the capability floor — which is the only configuration in which a productivity intervention actually scales.
The stack she deployed had four anchor components and a set of supporting integrations. Claude Code, the Anthropic-issued agent-coding environment, handles all Salesforce-native logic and the custom transforms that bridge ZoomInfo, the company's data lake on Snowflake, and the team's day-to-day tooling. Granola, the meeting-memory service, captures every customer conversation that the team participates in and produces structured outputs against templates Sokolowski wrote in March and revised in June. Linear, the project-execution platform, holds the workflow backlog and the workflow-improvement queue itself. Zapier, the lightweight automation layer, plumbs the events between systems where a more bespoke integration is not yet warranted. The configuration is unremarkable, which is the point. Helix Vertical did not adopt the leading edge of agent tooling. The team adopted the version of it that a competent revenue-operations leader could implement without engineering support — and then measured the result.
Lead enrichment: the baseline shift
Lead enrichment is the workflow where the agent gains are largest in absolute minutes and the most overstated in industry press. Helix Vertical's baseline numbers expose the gap. Before March 2025, the team ran an average of 280 enrichment passes per week across inbound and outbound leads. Each pass required between 11 and 19 minutes of human time — pulling ZoomInfo, cross-referencing LinkedIn, checking news search for company-context signals, scrubbing the entry against the team's ICP rubric, and logging into Salesforce. The mean was 14 minutes. Multiplied across the 280-pass week, that is 65 hours of revenue-operations capacity consumed by a workflow that produced no analytic output beyond a routed lead. Sokolowski's team had three full-time equivalents whose week was effectively this.
The agent workflow Sokolowski's team built around Claude Code took 11 weeks to converge to production quality. The first six weeks produced a workflow that ran in three minutes per pass but generated routing errors on 14 per cent of records — a rate the team judged unacceptable for a workflow that touched every inbound lead. The next five weeks were spent on the eval substrate: 1,400 historical enrichment passes were re-run through the agent, compared against the human-graded routing decision, and the divergences were used to refine the agent's account-fit rubric and the trigger logic for the manual-review queue. By the time the workflow cleared into production in early June 2025, the routing-error rate had fallen to 2.3 per cent, the per-pass time had risen marginally to 3.6 minutes to accommodate an additional verification step, and the workflow was passing 91 per cent of leads through to the routing engine without human intervention. The 9 per cent that fell to manual review were the cases where Salesforce held conflicting data on the account or where the agent flagged the source data as below confidence threshold. Both classes of escalation had been invisible in the pre-March baseline because the human reviewers had been processing them silently as routine.
The measured time saving is the figure Sokolowski has cited internally and that INTELAR is confirming here: from 14 minutes per pass to 3.6, across 280 passes per week, recovers 48.5 hours of team capacity. Three FTE-equivalent reduction. None of those three roles were terminated. Two were reallocated — one to a newly created account-research role that the team had not previously been able to staff, one to a dedicated renewal-forecasting analyst position. The third was a planned hire that was never made. The recovered capacity is the productivity gain. The reduction-in-force avoided is a count of one, not three, and Sokolowski insists on the distinction: agent workflows do not eliminate roles in a healthy operating environment. They eliminate the requirement to backfill them at a rate the company cannot afford.
Agent workflows do not eliminate roles in a healthy operating environment. They eliminate the requirement to backfill them at a rate the company cannot afford.
Account research and the templating discipline
Account research at Helix Vertical is the workflow where Sokolowski's team produced the most surprising figure. The pre-March 2025 baseline assumed account research ran at 47 minutes per account, averaged across the deep-dive briefs the team produced for opportunities above $150,000 in opening ARR potential. That figure had been benchmarked in October 2024 by a previous VP and was accepted as standard. When Sokolowski's team began instrumenting the workflow against an agent baseline in April 2025, the actual per-brief time was 71 minutes. The previous benchmark had measured the time a senior analyst took to produce a brief. The actual mean reflected the time the broader team needed, weighted by experience level. Variance had been compressing the benchmark downward. The instrumentation revealed the true cost.
The agent workflow that replaced the manual deep-dive operates against a templated rubric Sokolowski wrote in March and revised twice — once in May after the team's first 60 production briefs surfaced gaps in the regulatory-exposure section, once in August after the renewal-account version of the template proved insufficient for accounts with a multi-product footprint. The template specifies twelve sections: corporate overview, executive movements, financial trajectory, regulatory exposure, recent technology decisions, vendor stack inference, named champions, named blockers, peer-bench comparisons, expansion vectors, churn-risk indicators, and a freeform analyst-judgment section. The agent fills the first eleven sections from a defined source corpus that includes ZoomInfo, the team's own Salesforce instance, Granola call transcripts from any prior conversations with the account, publicly disclosed financial filings, and a curated set of trade-press sources Sokolowski's team validated against false-positive rates in March and April. The twelfth section remains human-written by design — the agent does not produce the analyst's judgment, only the substrate against which the judgment is formed.
The per-brief time fell from 71 minutes manual to 18 minutes assisted. Eleven minutes of agent runtime, seven minutes of human review and the analyst-judgment section. The team produces 38 deep-dive briefs per week on average. The recovered capacity is 33.6 hours per week — equivalent to roughly one full FTE — but Sokolowski uses this figure differently. Because the agent-assisted briefs are produced faster, the team is producing more of them. The pre-March 2025 baseline had been 22 briefs per week; the post-production rate is 38, with the additional briefs going to accounts in the $80,000 to $150,000 ARR band that had previously been below the briefing threshold. The result is not headcount reduction. It is coverage expansion at constant headcount, which is a different and harder-to-explain economic gain. The renewal forecasting workflow has benefited from the expansion: the company's net revenue retention forecast accuracy at 90 days improved from a ±9.4 per cent variance to a ±3.7 per cent variance over the four quarters of the measurement window. That improvement is, Sokolowski argues, downstream of the briefs going deeper into the mid-tier accounts where the previous coverage gap was hiding the renewal-risk signal.
Opportunity hygiene and the forecast loop
Opportunity hygiene is the workflow where Sokolowski's team learned the most about what agents cannot do. The pre-March 2025 baseline ran on a weekly Salesforce forecast call where the VP and the team reviewed every opportunity above $25,000 in expected close value. The call ran 90 minutes. The team committed to a forecast that, across the four quarters preceding the agent intervention, ran at a 13.8 per cent variance to actuals. Forecast misses skewed long: the team consistently over-committed and under-delivered, a pattern Sokolowski attributed to optimism bias compounded by the call's social dynamics. The reps wanted to commit. The VP wanted to commit. The numbers got committed.
The agent workflow Sokolowski's team built around Claude Code consumes every opportunity record above the $25,000 threshold, cross-references the activity log against the company's defined stage criteria, and produces a recommended stage classification that the rep can accept or override. The agent has access to every Granola transcript associated with the opportunity, every Salesforce activity log, and the company's calibrated stage-gate rubric. The first version of the workflow, deployed in late June 2025, produced a stage-classification accuracy of 78 per cent against the post-call human consensus — a figure Sokolowski's team initially considered acceptable. By August it was clear that the 22 per cent of cases where the agent and the human disagreed were not random. The agent was systematically more conservative than the team on opportunities where the only signal of progress was a recent customer-side scheduling email. The humans treated such signals as forward motion. The agent did not.
Sokolowski's team made a structural decision in September. Rather than tune the agent to match the human classification, they tuned the forecast call to weight the agent's classification at 60 per cent and the human classification at 40 per cent. The forecast call still happens, still runs 90 minutes, and still produces a committed forecast. The composition of the committed forecast has changed. Across the four quarters of the measurement window, the variance to actuals fell from 13.8 per cent to 4.9 per cent. The forecast accuracy gain is the headline number Helix Vertical's board has been quoting since Q1 2026. Sokolowski is more cautious. She notes that one full year of post-deployment data is insufficient to call the methodology validated. She also notes that the company's renewal book has been growing, which reduces forecast variance independently of the workflow change. The team will reach a real validation point at the close of Q4 2026, when the full four quarters of post-deployment data can be compared against a pre-deployment baseline that controls for renewal-book composition. Until then, the figure is a strong indicator. It is not yet a proof point.
Renewal forecasting and churn-risk scoring
The renewal-forecasting and churn-risk workflows ran in parallel from August 2025 onward and now produce the figure Sokolowski considers her most defensible. The combined workflow consumes the full customer-success and account-management activity log, the support-ticket queue, the product-usage telemetry that Helix Vertical's data lake aggregates from the deployed product, and a defined set of external signals — executive departures at the customer, regulatory enforcement actions in the customer's vertical, observable changes to the customer's procurement posture. The agent produces a weekly churn-risk score for every installed-base account, classifies the score against four bands, and routes the high-risk band to a dedicated customer-success motion that did not exist before the workflow was deployed.
The measured outcome is contained in the company's ARR-coverage ratio — the ratio of installed-base ARR that is actively managed under a defined customer-success motion to the total installed-base ARR. In March 2025, the ratio was 64 per cent. The company's top 40 accounts by ARR received a full customer-success motion; the next 80 received a touch motion; everything below received an automated outreach motion that was, in practice, untracked. By April 2026, the ratio had moved to 91 per cent. The improvement did not come from headcount: the customer-success function added two roles during the window, well below the pace required to cover the additional accounts at the pre-March staffing ratio. The improvement came from the agent's ability to route accounts into the appropriate motion based on the churn-risk score rather than the legacy ARR-tier classification. A $40,000 ARR account at high churn risk now receives the full motion. A $400,000 ARR account at low churn risk does not consume a dedicated CS role that another account could be using.
The churn-risk scoring workflow has produced its share of failure modes, and Sokolowski's team has logged them with the same discipline they applied to the lead-enrichment work. The agent has surfaced false positives at a rate of 11 per cent — accounts flagged as high-risk that the customer-success team, after manual review, judged to be in normal operating state. The team treats these false positives as inputs to the rubric refinement rather than as agent errors, on the theory that a workflow that is right 89 per cent of the time and produces 11 per cent excessive outreach is preferable to a workflow that is right at the population mean but misses the specific accounts where outreach would have prevented churn. The asymmetry of the cost — outreach is cheap, churn is expensive — favours the over-flagging bias. The agent is configured to be cautious. The team is configured to filter.
Failure modes and the instrumentation discipline
The Helix Vertical programme has not been frictionless, and the value of the field study is partly in the failures the team has logged. The first significant failure mode was the lead-enrichment routing error in the initial six weeks of deployment, when the agent's account-fit rubric was misclassifying 14 per cent of inbound leads. The cost of that error was relatively low because the manual-review queue caught the misroutes before they generated downstream pipeline damage. The team's response — eleven weeks of eval substrate work before clearing the workflow into production — set the operating discipline for every subsequent workflow. No workflow goes live without a defined eval set, a documented threshold, and a designated review cadence.
The second failure mode was more expensive. In late August 2025, the renewal-forecasting workflow produced a forecast for a Q4 2025 cohort that materially underestimated the renewal probability for a specific class of multi-product accounts. The agent was treating the multi-product footprint as a single account-level signal rather than a composite of product-level signals, and the under-weighted product-level signals were where the renewal momentum was strongest. The forecast missed the cohort's actuals by 8.3 per cent. The team caught the error in the Q4 2025 close review and traced it to a rubric definition that had been written when most accounts were single-product. The rubric was rewritten in January 2026, the workflow re-ran the four prior quarters of data with the updated rubric, and the team produced a calibrated baseline that the board accepted as the new reference. The lesson Sokolowski drew was procedural: workflows that were written when the company's product footprint was simpler need an explicit re-validation step every time the product footprint expands materially.
The third failure mode is the one Sokolowski is least confident the team has solved. The opportunity-hygiene workflow, after the September weighting change, has produced a forecast that the reps trust less than the pre-March process. The reps used to argue the forecast in the call. They now accept the agent's classification more often, partly because the data behind it is more comprehensive than any individual rep can produce in the call. The participation pattern in the forecast call has shifted: the call runs 90 minutes, but the meaningful conversation about specific deals has shrunk from roughly 60 minutes to roughly 35. The reps spend less of the call defending their classifications and more of the call discussing the deals where the agent classification and the rep's intuition diverge. The forecast accuracy is better. The ownership of the forecast is more ambiguous. Sokolowski's team is running a working group through Q2 2026 to determine whether the ambiguity is a transitional artefact or a structural feature of agent-weighted forecasting, and the answer matters for every revenue-operations team that follows Helix Vertical down this path. The forecast is the team's commercial commitment. If the commitment is ambiguously owned, the team is in unfamiliar terrain.
What to watch
Helix Vertical's twelve-month measurement window is now closed, but the operating questions that the programme has raised will not resolve cleanly until the company's growth profile changes again. The next decision window arrives in Q3 2026 when the team plans to publish an internal review of the methodology and decide whether the agent stack scales to a 60-person operations team at the company's projected 2027 bookings profile.
- Whether the forecast-ownership question resolves before the team's Q3 2026 review; if the reps continue to defer to the agent's classification more than Sokolowski thinks is healthy, the working group will recommend a structural change to the call format that could move the agent's role from a default classification to a tie-breaker only, which is a meaningfully different operating posture.
- Whether the multi-product rubric Sokolowski's team wrote in January 2026 holds against the expanded product footprint the company plans to ship in H2 2026; the renewal-forecasting workflow has been re-validated once, but the cadence at which the team will need to re-validate is not yet established, and an annual re-validation is more expensive than a quarterly one.
- Whether the ARR-coverage ratio improvement persists when the installed base grows materially; the team has moved from 64 per cent to 91 per cent coverage on a base that grew about 30 per cent during the window, and the question is whether the agent's churn-risk routing can hold the coverage ratio if the base grows another 50 per cent without proportional CS hiring.
- Whether the eleven per cent false-positive rate on churn-risk scoring becomes a customer-side complaint vector; the team is not yet hearing about excessive outreach from the installed base, but a quarter of high-touch outreach to accounts that did not need it is a real customer experience cost that the team is monitoring against survey data.
- Whether Sokolowski's team publishes the workflow rubrics externally, as the team's investors have begun encouraging; the rubrics are the documented form of the operating discipline, and the question of whether they constitute proprietary intellectual property or industry-shareable methodology is itself the question that Helix Vertical's board will have to answer in the next two quarters.
Frequently asked
- What is the agent stack Helix Vertical's revenue-operations team is running?
- Four anchor components plus integrations. Claude Code handles the Salesforce-native logic and the custom transforms across ZoomInfo, the company's Snowflake data lake, and the day-to-day tooling. Granola captures every customer conversation the team participates in and produces structured outputs against templates Sokolowski wrote. Linear holds the workflow backlog and the workflow-improvement queue. Zapier plumbs the events between systems where a more bespoke integration is not yet warranted. The stack is intentionally unremarkable. The investment is in the rubrics and the eval substrate, not in the leading edge of agent tooling.
- How much time did the lead-enrichment workflow actually save, and how should that figure be interpreted?
- The workflow moved per-pass time from 14 minutes to 3.6 minutes, across 280 passes per week, which recovers 48.5 hours of team capacity. The figure should not be interpreted as headcount eliminated. Two roles were reallocated to functions the team had not previously been able to staff — a dedicated account-research role and a renewal-forecasting analyst position — and one planned hire was never made. The reduction-in-force avoided count is one. The productivity gain is the work the team can now do that it could not previously.
- Why did the opportunity-hygiene workflow not eliminate the weekly forecast call?
- Because the forecast call serves a function the agent does not replace. The agent produces a stage classification with high accuracy and comprehensive data. The call produces the team's commercial commitment to a forecast and the conversation about deals where the agent and the rep disagree. Sokolowski's team treats the agent classification as a 60 per cent input to the call's committed forecast, not as a replacement for the call. The shift in the call's substance — less defence of classifications, more analysis of divergence — is itself the productivity gain, and it does not come from eliminating the call.
- How does the churn-risk workflow handle the false-positive rate?
- The agent surfaces high-risk flags at a rate that includes about 11 per cent false positives — accounts the customer-success team, after manual review, judged to be in normal operating state. The team treats the false-positive rate as a feature of the rubric rather than a bug, on the theory that the asymmetry of the cost favours over-flagging. Outreach is cheap. Churn is expensive. The agent is configured to be cautious; the team is configured to filter. The team monitors the false-positive rate against survey data from the installed base and will adjust the rubric if the rate generates customer-experience complaints. It has not yet.
- What is the ARR-coverage ratio, and why did Sokolowski's team treat it as the headline metric?
- The ARR-coverage ratio is the share of installed-base ARR actively managed under a defined customer-success motion, divided by the total installed-base ARR. It moved from 64 per cent in March 2025 to 91 per cent in April 2026. Sokolowski's team treated it as the headline metric because it captures the workflow gain without reducing it to a headcount number. The improvement came from the agent's ability to route accounts based on churn-risk score rather than ARR tier — which means smaller accounts at high risk get the full customer-success motion they need, and larger accounts at low risk no longer consume a dedicated role that another account could use. The ratio is a coverage figure, not a productivity figure. It is also the figure the company's investors find easiest to translate into a renewal-rate forecast.
- Is this methodology portable to other revenue-operations teams?
- In principle yes, but the carry-over is not the tools. The carry-over is the discipline. The Helix Vertical programme worked because Sokolowski's team built rubrics, instrumented baselines, defined eval substrates before deploying any workflow into production, and treated failure modes as inputs to rubric refinement rather than as deployment failures. A team that adopts Claude Code, Granola, Linear, and Zapier without the rubric discipline will get a different result. The stack is the cheaper half of the work. The operating method is the expensive half.
Twelve months of measured data on a 40-person revenue-operations team running a defined agent stack produces a pattern that the productivity literature has been claiming without evidence for at least two cycles. The pattern is real. The headcount that did not get added is the productivity gain. The reduction-in-force avoided count is the metric the board cares about. The ARR-coverage ratio is the metric the investors cite. The forecast accuracy improvement is the metric Sokolowski's team is least willing to call definitive, because one year of data on a growing renewal book does not fully control for the variance the agent intervention is supposed to explain. The methodology will be re-validated at the close of Q4 2026, and the answer that emerges will either confirm Helix Vertical's programme as a portable reference or expose a confound the team's instrumentation has not yet caught.
Sokolowski's argument — that the value of an agent stack is downstream of the rubric discipline the team applies to it — is the load-bearing claim in the Helix Vertical study. The tools matter. They are not the leading edge of the available stack. They are the version of the stack that a competent operations leader could deploy in a year without engineering support. The rubrics are the work. The eval substrates are the work. The failure-mode logs are the work. The boards that funded the programmes that produced last year's productivity press tend to skip past this paragraph, and that is a forecast more reliable than any agent in this article: the gap between the companies that read this and adopt the discipline and the companies that read this and adopt only the tools will be visible by Q3 2027. INTELAR will be tracking both.
More from Productivity →