Chapter 4

Operating expense: the fourteen levers from engineering to the executive team

The most levers, the smallest individual moves, and the one benchmark that sets the ceiling: the best-run companies closed this gap before any of these tools existed.

24 min read

Chapter 4 of 10Operating expense24 min

Hackett's July 2026 benchmark of the thousand largest North American public companies puts median SG&A at 16.2% of revenue and the first quartile at 7.8%.1 That 8.4-point gap is the size of the operating-expense prize, it predates the technology, and it reframes every card in this chapter: these levers are the cheapest way yet to reach a level the best companies already reached by other means. The research and development levers are different in kind, because engineering is where the technology's task-level evidence is most mixed and where the downstream costs (defects, review load, incidents) are best measured.

Three cautions apply across the chapter. Headcount falls through attrition and role redesign, never in the first two quarters, when retention through the ownership transition matters most. The transformation office owns measurement and never a P&L line; the function lead owns the target and carries it in compensation. And several widely quoted benchmarks in this area do not exist: the "Gartner 20 to 30% ticket deflection" figure appears in no Gartner document, and vendor deflection percentages of 65% and 75% are absent from the pages that supposedly carry them.2

O1 · Engineering throughput

What it is. Coding agents on every task with human review as the binding constraint, test generation bringing legacy coverage to a defensible level, a review agent doing first-pass review, triage agents routing defects with a proposed fix, and documentation agents keeping runbooks and references current from the code. Adoption is mandated and measured; the engineering leader owns the review-capacity plan; change failure rate is a gate.

The line it moves. Research and development; and support and retention through defect rates.

Typical range and the evidence. Strong, and mixed, and the honest reading is that it depends on who is measured and what happens downstream. R&D from the low twenties to the high teens as a share of revenue while shipping three times the feature volume is the case-study range. METR's randomized trial of sixteen experienced open-source developers on their own repositories found they "take 19% longer to complete issues" with AI tools while believing afterwards that "AI had sped them up by 20%"; its February 2026 follow-up estimated 18% faster for the ten who returned, with a confidence interval from 38% faster to 9% slower, and then called the design "likely a bad proxy."3 Cui and colleagues' randomized rollout to 4,867 developers at three companies found "26.08% (SE: 10.3%)" more completed tasks.4 Downstream is what matters to an owner: Faros AI's 2026 telemetry across 22,000 developers found task throughput up 33.7% alongside 54% more bugs per developer, a fivefold increase in median review time and 242.7% more incidents per pull request;5 Google's 2025 DORA report found AI adoption positively related to throughput and "a negative relationship with software delivery stability."6 Company-level evidence is clearer than task-level: High Alpha's 2025 benchmarks found 69% of software companies between $5M and $20M of ARR had reduced headcount because of AI, engineering most of all, and Vista reports R&D across 54 selected companies falling "from 21.9% in 2023 to 20.1% in 2024 and 19.2% in 2025."7 The maintainability data from P1 (duplication up 81%, refactoring down 70%) is the cost of doing this without the gates.8

Conditions. Modern source control and CI with the test harness in place; an agentic coding platform with enterprise controls; a QA discipline that treats agent-written code as untrusted; security scanning mandatory on generated code; synthetic or de-identified test data; leadership that holds quality metrics while velocity rises.

Kill criteria. Change failure rate or incident count above baseline for a month pauses the expansion of agent scope. No agent merges to production, ever.

  • Meaningful releases per year
    Revenue (indirect)
    At close2 to 4
    Year 16
    Year 310
  • Automated test coverage on core modules
    Quality, R&D
    At closeLow
    Year 150%
    Year 375%
  • Change failure rate
    Support, retention
    At closeBaseline
    Year 1−30%
    Year 3−50%
  • Engineering hours on maintenance
    R&D
    At close~40%
    Year 130%
    Year 320%
  • R&D, % of revenue
    EBITDA
    At close20 to 25%
    Year 1−2 points
    Year 3−5 points

Horizon and stage. Deploy, then Reshape as the mix shifts senior. Day one hundred to year two.

How it translates. A: a twelve-year-old codebase, so test generation comes first. B: the rules engine must be testable before agents touch it. C: mobile and payments code, where the security gate is non-negotiable. D: the largest team and the largest AI-generated share already; the lever is the gates, not the tools. E: several codebases; O2 and X4 decide which ones survive.

O2 · Tech-debt paydown and migrations

What it is. Framework, language, dependency and data-model migrations that were uneconomic at human cost, run by agent under engineer supervision: the Java upgrade, the framework move, the monolith split, the acquired codebase folded into the platform.

The line it moves. R&D, as a one-time cost that is a fraction of the historical one; and it enables G4 (a modern stack hosts cheaper) and X4 (add-ons consolidate).

Typical range and the evidence. Strong, at scales no small company will see, with a mechanism that transfers. Google's January 2025 paper on internal migrations reports that in one large migration "80% of the code modifications in the landed CLs were fully AI-authored," with about a 50% reduction in time, and that in a test-framework migration "~87% of the code generated by AI ended up committed without any change" across 5,359 files in three months; the bottleneck moved to reviewer capacity.9 Amazon's account of its Java upgrade campaign cites "over 4,500 years of development work" saved and "$260 million dollars in annual cost savings," and the two figures measure different things, the first a modeled labor estimate and the second a realized infrastructure saving from newer runtimes; they are routinely conflated and should not be.10

Conditions. A migration plan with the target state written down; senior reviewers with the time to review, which is the scarce resource; a rollback path; the test coverage from O1 first.

Kill criteria. A migration whose reviewer hours exceed the plan by 50% at the halfway point is paused and rescoped.

  • Migrations completed that were previously deferred
    R&D
    At close0
    Year 11 to 2
    Year 33 to 5
  • Cost per migration vs historical estimate
    R&D
    At closeBaseline
    Year 1−40%
    Year 3−60%
  • Reviewer hours as share of migration effort
    R&D
    At closeMinor
    Year 1Majority
    Year 3Majority
  • Legacy modules with test coverage above 50%
    Quality
    At closeFew
    Year 1Half
    Year 3Most

Horizon and stage. Reshape. Years one to two.

How it translates. A: the front end is modern and the back end is not; one migration carries the plan. B: the rules engine is rewritten to be testable. C: the payments integration is migrated to the chosen processor's platform. D: continuous; the largest program. E: the acquired codebases, which is X4.

O3 · Product management and research

What it is. Discovery, specification, analysis and documentation work compressed by agents that synthesize support conversations, usage data and interviews into ranked opportunity briefs, draft specifications and generate acceptance criteria, so that the product-manager-to-engineer ratio changes and the discovery cycle shortens.

The line it moves. Research and development.

Typical range and the evidence. Weak, and kept for that reason. Practitioner accounts describe the compression; no independent study or named-customer figure measures it. The card exists because the ratio is a real line in the R&D budget and because an owner who changes it should measure what happens to the release quality and the roadmap hit rate.

Conditions. The voice-of-customer digest from R5 and G6; usage data at the feature level; a product leader willing to run the function as a measured experiment.

Kill criteria. Roadmap items shipped on time falling below baseline for two quarters after the ratio changes.

  • Engineers per product manager
    R&D
    At closeBaseline
    Year 1+25%
    Year 3+50%
  • Discovery cycle, idea to specification
    R&D
    At closeWeeks
    Year 1Days
    Year 3Days
  • Roadmap items shipped on time
    Quality
    At closeBaseline
    Year 1Held
    Year 3Held or better

Horizon and stage. Reshape. Year one.

How it translates. Most material for D, where the product organization is large, and B, where the research is regulatory. Marginal elsewhere.

O4 · Sales productivity

What it is. Fewer, better-supported sellers: research agents assembling the target universe and enriching contacts, agent-drafted and human-approved outbound, call intelligence, proposal drafting, CRM hygiene done by agent, forecasting from engagement evidence, and the sales-development function largely automated with a person approving what goes out.

The line it moves. Sales and marketing expense, with the output metrics in R6 and R7.

Typical range and the evidence. Medium. Sales spend rising in absolute terms while pipeline per representative doubles is the case-study shape; the efficiency evidence is Vista's, with sales and marketing spend across 55 selected companies falling "from 30.4% in 2023 to 27.6% in 2024 and 25.6% in 2025."7 Momentum quotes Ramp's head of go-to-market systems saying the tool "has cut the time in half for sellers to progress their deals in Salesforce" and titles its ScyllaDB case with a 30% productivity gain; Clay's A-LIGN case reports an "83% reduction in research costs."11 All vendor, all named. The independent evidence on AI-written outbound (1.4% positive replies against 2% to 4% human) is in R7 and is the reason the person stays in the loop.

Conditions. R6's CRM and tooling; compensation on representative productivity rather than headcount; per-domain volume caps; the founder's transition plan, because founder-led selling is the usual baseline.

Kill criteria. Cost per qualified meeting not down 30% in three quarters; or any sequence sent without human approval in the first two quarters.

  • Qualified meetings per month
    Revenue
    At closeBaseline
    Year 1+80%
    Year 3+180%
  • Cost per qualified meeting
    Sales
    At closeBaseline
    Year 1−40%
    Year 3−55%
  • New-logo ARR per seller
    Sales efficiency
    At closeBaseline
    Year 1+40%
    Year 3+100%
  • Sales and marketing, % of revenue
    EBITDA
    At closeBaseline
    Year 1Flat to +1 point
    Year 3Flat, on a larger base

Horizon and stage. Reshape. Years one to two.

How it translates. A: removing founder dependence. B: enterprise sellers supported on regulatory research. C: velocity sales with the demo environment and the payments pitch. D: sales-assist on product signals. E: enterprise cycles where the proposal and the questionnaire are the work.

O5 · Marketing production

What it is. Content, creative and campaign operations produced by agents with expert review at a fraction of the agency and headcount cost; the same team, three times the output; the measured outcome in R7.

The line it moves. Sales and marketing expense (agency and contractor spend, and headcount held flat while output rises).

Typical range and the evidence. Medium. Jasper's named results (WalkMe "3,000+ hours saved in content creation time," Akbank "40% reduction in time spent creating content") measure production time; nothing independent measures the pipeline effect.12 The card's value is the agency line, which is real and visible.

Conditions. Brand and claims guidelines the drafting agents enforce; expert review before any regulatory content is published; no generated testimonials, ever.

Kill criteria. Agency and contractor spend not down 40% after two quarters with output held.

  • Content pieces per month, expert-reviewed
    Marketing
    At close2
    Year 112
    Year 320
  • Agency and contractor spend
    Marketing
    At closeBaseline
    Year 1−40%
    Year 3−60%
  • Marketing headcount
    Marketing
    At closeBaseline
    Year 1Flat
    Year 3Flat
  • Event follow-up complete within 48 hours
    Revenue
    At closeRarely
    Year 1Always
    Year 3Always

Horizon and stage. Deploy. Day one hundred to year one.

How it translates. Roughly uniform; largest in D, where the content engine is the funnel.

O6 · Finance close and FP&A

What it is. Close automation (reconciliations, accrual proposals, revenue-recognition schedules, flux analysis with explanations, checklists), a driver-based forecast refreshed weekly from the warehouse, board and investor packages generated from the KPI tree with narrative drafted for the chief executive to edit, and audit and tax requests answered from a maintained evidence library. The controller reviews exceptions and becomes a finance lead supported by agents.

The line it moves. General and administrative expense; planning quality.

Typical range and the evidence. Strong for the benchmark, medium for the cases; finance has the best independent evidence in the back office. Days to close from the low teens to four or five, and finance headcount halved, is the case-study range. APQC's benchmarks put the monthly close median at six days with top performers at five or fewer.13 Choi of MIT Sloan and Xie of Stanford studied 79 small and medium companies on an AI-enabled accounting platform and found accountants could cut 7.5 days off the monthly close and support 55% more clients per week.14 Numeric's November 2025 announcement says Brex "saw their match rate jump from 30% to over 90%" on reconciliations.15 Pigment's Supercell story, in the customer's own words, is that it "took 2 days for 4 people to update our large P&L spreadsheet" and afterwards "those same updates took me 4 minutes."16 MIT NANDA's finding that "some of the most dramatic cost savings we documented came from back-office automation," with faster payback than the sales and marketing programs that absorbed about 70% of budgets, is the reason this lever is deployed first.17

Conditions. A modern general ledger and billing system (replace anything that cannot integrate); the warehouse as the single source for revenue and usage; close and FP&A platforms purchased and connected; a controller willing to redesign the function around exceptions; the auditor's review of revenue-recognition logic before it runs unattended.

Kill criteria. Close not under eight days by month nine. Segregation of duties is preserved throughout: agents propose, humans approve journal entries above thresholds.

  • Days to close
    G&A
    At close10 to 15
    Year 16
    Year 34
  • Forecast error, quarterly revenue
    Planning
    At close±8%
    Year 1±4%
    Year 3±2%
  • Hours to produce the board package
    G&A
    At close40
    Year 110
    Year 34
  • Finance headcount
    G&A
    At closeBaseline
    Year 1−25%
    Year 3−50%

Horizon and stage. Deploy. Day one hundred to year one.

How it translates. Nearly uniform; the absolute saving scales with the size of the finance team, so it is largest in E and D, and the proportional saving is largest in A, where four people become two.

O7 · Accounts payable and expense

What it is. Invoice capture, coding, routing, approval and scheduling by agent; expense capture and coding; vendor onboarding with the compliance checks built in.

The line it moves. General and administrative expense.

Typical range and the evidence. Medium. No-touch rates of 70% to 85% at maturity. Vic.ai says clients "can achieve up to an 85% no-touch invoice rate" with its named cases landing at 72% to 78%; Brex reports "over 65% of all expenses on Brex are fully automated" with more than 90% of AI coding suggestions accepted; Ramp reports 3.5 times more auto-coding with its accounting agent.18 All vendor, most named.

Conditions. The general ledger from O6; a vendor master that is clean; approval thresholds written down; payment release always human.

Kill criteria. No-touch rate below 50% after two quarters.

  • AP and expense no-touch rate
    G&A
    At close~10%
    Year 170%
    Year 385%
  • Days from invoice receipt to approval
    G&A, vendor terms
    At closeWeeks
    Year 1Days
    Year 3Day
  • Duplicate or erroneous payments
    G&A
    At closeUnknown
    Year 1Measured
    Year 3Near zero

Horizon and stage. Deploy. Day one hundred.

How it translates. Uniform and small; a purchase configured by the controller.

O8 · Legal and contracting

What it is. First-pass review and redlining of customer, vendor and partner agreements against a written playbook; extraction of every executed agreement into a structured database that feeds renewals (R4) and pricing (R1); outside-counsel triage so that counsel sees non-standard terms, indemnities, liability caps and regulated-data language and nothing else.

The line it moves. General and administrative expense (outside counsel); sales cycle through contract turnaround.

Typical range and the evidence. Strong for the productivity effect, medium for the cases, and weak for the claim most often made. Contract turnaround from two weeks to days is well supported: Ironclad reports Poshmark "cutting their contract turnaround times down by 96%" and Hormel moving "from 12 weeks down to consistently 3 weeks"; Ironclad's own benchmark puts the average time to execute at 42 days with 60% of contracts on counterparty paper.19 Harvey's named in-house results are in P2. The first randomized controlled trial of legal AI, by Schwarcz and colleagues, found "statistically significant gains of anywhere from 50% to 130%" in productivity on five of six tasks, with the caveat that the subjects were law students on academic exercises.20 What the evidence does not support is that AI cuts outside-counsel spend: CLOC's 2026 survey of 135 departments found the share expecting to increase outside-counsel spend fell from 58% to 37%, a cooling of expected increases rather than a decline, and Thomson Reuters' 2026 report found 36% of general counsel expecting to increase it against 20% planning to decrease.21 Thomson Reuters' Future of Professionals survey of 1,816 professionals found 74% using AI several times a week and 91% saying their organizations fall short of what it could deliver; with a named AI strategy, 66% said it met expectations, and without one, 22%.22 The case-study target of outside counsel down 40% is therefore an internal target at a company whose counsel spend is mostly routine paper, and it is graded as such.

Conditions. A contract lifecycle platform with AI review; a written playbook of fallback positions and non-negotiables; the workflow designed with counsel so that privilege is preserved where counsel directs it.

Kill criteria. Contract turnaround not under five days by month nine on standard paper.

  • Contract turnaround, standard paper
    Sales cycle, G&A
    At close14 days
    Year 15 days
    Year 33 days
  • Agreements in a structured database
    Renewals, pricing
    At close0%
    Year 1100%
    Year 3100%
  • Outside-counsel spend on routine paper
    G&A
    At closeBaseline
    Year 1−25%
    Year 3−40%
  • Share of contracts on the company's paper
    Risk
    At closeMinority
    Year 1Rising
    Year 3Majority

Horizon and stage. Deploy. Year one.

How it translates. A: customer paper and business associate agreements. B: regulated-data terms in every contract; the lever is also product evidence. C: merchant agreements at volume, where the database matters most. D: uniform terms of service; small. E: bespoke enterprise paper; the largest absolute saving.

O9 · Compliance operations

What it is. Continuous control monitoring and evidence collection for SOC 2, ISO 27001, HIPAA and whatever else the vertical requires, through a purchased platform connected to cloud, identity, HR and device systems; security questionnaires answered from the evidence library; policies drafted, versioned and attested through the same platform. It is also a sales lever (R6) and a posture lever (X1).

The line it moves. General and administrative expense; sales cycle.

Typical range and the evidence. Weak, and the vendors' own data is the reason. Vanta's claim of "automating 85% of evidence collection" carries no citation; IDC's Vanta-sponsored snapshot reports "82% less staff time needed per framework and attestation-related audit" on an undisclosed sample; Forrester's Drata study, commissioned by Drata, models audit preparation falling from 980 to 220 hours a year and framework maintenance down 76%, with most of the modeled benefit in avoided consulting fees.23 Vanta's own State of Trust 2025, from 3,500 respondents, reports twelve working weeks a year spent on compliance tasks, up from eleven, and nine weeks on vendor security reviews, up from seven: the burden is growing faster than the tooling, by the automation vendor's own account.24 Samsara's consolidation of "820 controls across 10 frameworks" into about 260 with an agent is the one named enterprise figure.25 No independent evidence shows certification lifting win rates, and the defensible claim is that it removes a gate and compresses the questionnaire step.

Conditions. Cloud, identity, HR and device systems the platform can connect to; a named owner; the annual risk analysis that HIPAA requires and that the platform feeds rather than replaces.

Kill criteria. Evidence collected automatically below 50% after two quarters, or a questionnaire answer wrong on a control that matters.

  • Compliance evidence collected automatically
    G&A
    At close~10%
    Year 170%
    Year 385%
  • Security questionnaire turnaround
    Sales cycle
    At close10 days
    Year 12 days
    Year 31 day
  • Frameworks maintained per compliance FTE
    G&A
    At close1
    Year 12
    Year 33
  • Time to add a framework
    G&A, sales
    At closeQuarters
    Year 1Months
    Year 3Weeks

Horizon and stage. Deploy. Day one hundred to year one.

How it translates. A: HIPAA and SOC 2, and the risk analysis is the document the regulator asks for. B: the company's own compliance is part of its product credibility; the evidence library is customer-facing. C: PCI scope, decided in P7. D: SOC 2 and ISO at scale; the questionnaire volume is the cost. E: enterprise customers' questionnaires and, in Europe, DORA obligations as an ICT third-party provider.26

O10 · Recruiting and people operations

What it is. Job descriptions, sourcing, structured screening, scheduling and interview summaries by agent; an employee helpdesk answering policy, benefits and payroll questions; onboarding plans and the AI-skills curriculum; the redeployment program that moves people from automated work to growth work, which is what makes the headcount plan humane and achievable.

The line it moves. General and administrative expense; manager time; time to hire as a growth input.

Typical range and the evidence. Medium. Paradox's 7-Eleven case reports time to hire falling "from over 10 days in most cases" to "under 5 days" in high-volume store hiring; knowledge-worker evidence is thinner, and the case-study target is a 30% to 50% cycle-time reduction.27 Workday's September 2026 update reports more than 5,500 customers using its agents, with named outcomes at Mohegan (configuration analysis from hours to thirty seconds) and CLEAResult.28 No credible study measures AI's effect on quality of hire, and the card does not claim one. The legal exposure is measured: in Mobley v. Workday a federal court certified in May 2025 a nationwide collective of applicants aged forty and over against the software vendor on an agency theory, which is why screening uses structured criteria, is audited for adverse impact quarterly, and never makes the decision.29

Conditions. An HR platform with helpdesk and workflow capability; an applicant tracking system; a written redeployment policy; manager training on leading teams that include agents.

Kill criteria. Any adverse-impact audit finding pauses screening automation until resolved.

  • Time to hire (days)
    G&A, growth
    At closeBaseline
    Year 1−35%
    Year 3−50%
  • Employee questions answered without HR
    G&A, manager time
    At close0%
    Year 150%
    Year 370%
  • Employees through the AI-skills curriculum
    Adoption
    At close0%
    Year 1100%
    Year 3100%, annual refresh
  • Roles redeployed vs eliminated
    All cost lines
    At closen/a
    Year 1Majority redeployed
    Year 3Per plan

Horizon and stage. Deploy. Year one.

How it translates. Uniform; the redeployment program matters most where the headcount change is largest (A, C, E).

O11 · IT service desk

What it is. Employee support, provisioning, deprovisioning and access requests resolved by agent with approval workflows; alert correlation, runbook execution and post-incident drafting for the operations team.

The line it moves. General and administrative expense; incident resolution time as a retention input.

Typical range and the evidence. Medium, with the benchmark caution above. Moveworks reports Databricks' employee assistant climbing from 10% adoption at launch to 73%, "with about half of inquiries handled entirely by the bot."30 SolarWinds' analysis of more than 60,000 ITSM records found generative-AI users resolving tickets 17.8% faster on average, with the top adopters at 54.3%.31 Microsoft's observational study of 95,522 security incidents found generative AI "associated with a 30.13% reduction in security incident mean time to resolution," and its randomized trials found professionals 22% faster and 7% more accurate.32 The platform incumbent's $2.85B purchase of Moveworks in March 2025 is the market's view that this is a feature of the service platform rather than a product.33

Conditions. Centralized logging and observability; documented runbooks; an approval step on any production change for six months.

Kill criteria. Tickets resolved without a person below 30% after two quarters.

  • Employee IT requests resolved by agent
    G&A
    At close0%
    Year 140%
    Year 360%
  • Mean time to resolve incidents
    Retention, R&D
    At closeBaseline
    Year 1−40%
    Year 3−50%
  • IT, security and DevOps headcount
    G&A
    At closeBaseline
    Year 1Flat
    Year 3−25 to −33%

Horizon and stage. Deploy. Day one hundred.

How it translates. Material for D and E, where the desks are large; a feature of the HR and operations platforms for A, B and C.

O12 · Procurement and SaaS spend

What it is. License reclamation, renewal negotiation and intake-to-pay through a spend platform, with the SaaS portfolio managed as a line rather than as a series of card charges.

The line it moves. General and administrative expense (and R&D and S&M where the licenses sit).

Typical range and the evidence. Strong for the waste, medium for the savings. Zylo's 2026 index of more than 40 million licenses and $75B under management found 36% of licenses unused and median SaaS spend of $9,455 per employee, with 78% of IT leaders reporting unexpected charges tied to consumption or AI pricing.34 Forrester's Zip study, commissioned by Zip, reports "3.3% average savings on all spend flowing through the platform" and a 70% reduction in request cycle time on a composite of four very large enterprises; use the 3.3%, not the 386% return-on-investment headline.35 Hackett's 2026 procurement survey found 69% of organizations accessing AI through capabilities embedded in existing platforms, which argues against standalone tools at this scale.36

Conditions. A complete inventory of subscriptions, including the ones on expense cards; a renewal calendar; someone who owns the number.

Kill criteria. SaaS spend per employee not down 15% after the first renewal cycle.

  • Unused licenses, % of total
    G&A
    At close~36%
    Year 115%
    Year 3Under 10%
  • SaaS spend per employee
    G&A
    At closeBaseline
    Year 1−20%
    Year 3−30%
  • Renewals negotiated with usage data
    G&A
    At closeFew
    Year 1All
    Year 3All

Horizon and stage. Deploy. Day one hundred.

How it translates. Small in absolute terms for A and B, large for D and E, and the first thing to do in an add-on (X4).

O13 · The executive operating system

What it is. The weekly business review generated from the warehouse every Monday, a live KPI tree from enterprise value down to agent-level metrics, decisions logged with owners and dates, the register of every production agent with its owner, scope, evaluation status, cost and measured value, and the monthly value review at which initiatives are judged against their baselines and killed on their dates. It is the mechanism that makes the rest of the book visible and accountable.

The line it moves. General and administrative expense (executive and analyst time) and, through governance, every other line.

Typical range and the evidence. Strong for the mechanism, because the evidence is about what happens without it. FTI's finding that 66% of PE leaders see benefits inside twelve months while only 31% call implementation efficient, and its companion finding that 95% of funds meet their business case while 7% of companies run AI at enterprise scale, describe programs without a register.37 Gartner's forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027 should be read as a description of hygiene: a company that has cancelled six initiatives on schedule and kept eight that moved their lines is running the plan.38 NIST's Measure function, "quantitative, qualitative, or mixed-method tools, techniques, and methodologies to analyze, assess, benchmark, and monitor AI risk and related impacts," is the standard the register implements for value as well as risk.39 Thomson Reuters' 66%-against-22% split by whether the organization has a named AI strategy is the closest thing to a causal claim that governance, not tooling, decides the realized value.22

Conditions. The warehouse and KPI tree from P7; a transformation office of two or three that owns measurement and no P&L line; a cadence the chief executive enforces.

Kill criteria. This is the lever that kills the others. Its own failure mode is an initiative "in progress" at month eighteen with no measured movement.

  • Hours to produce the weekly business review
    G&A
    At close8
    Year 11
    Year 30.5
  • Initiatives with measured value vs plan
    Value tracking
    At closen/a
    Year 1100%
    Year 3100%
  • Initiatives killed on their kill date
    Governance
    At closen/a
    Year 1Some
    Year 3Some, every year
  • Agents in the register with passing evaluations
    Risk
    At closen/a
    Year 1100%
    Year 3100%

Horizon and stage. Reshape. Day one hundred to exit.

How it translates. Identical in every company. The register is also the data room at exit (X6).

O14 · G&A benchmark closure

What it is. The sum of O6 through O12, stated as one target against the external benchmark so that the board can see the whole rather than the parts.

The line it moves. General and administrative expense as a share of revenue.

Typical range and the evidence. Strong. Hackett's 16.2% median against 7.8% first quartile is for large public companies and the shares differ at $10M, but the shape holds: the case study takes G&A from 15% to 10% of revenue, and SaaS Capital's 2026 medians put G&A at 15% of ARR for private companies in the $3M to $5M band.40 Five points of G&A is the target for a small company, three for a large one, and the levers above are how.

Conditions. Every O lever with its own baseline; a finance lead who reports the sum monthly.

Kill criteria. None; the parts carry their own.

  • G&A, % of revenue
    EBITDA
    At close12 to 15%
    Year 1−2 points
    Year 3−4 to −5 points
  • Revenue per G&A employee
    G&A
    At closeBaseline
    Year 1+30%
    Year 3+80%

Horizon and stage. Reshape. Years one to three.

How it translates. Uniform in shape; the case study shows the arithmetic.

Sources

  1. The Hackett Group, "SG&A Costs Reach Five-Year Highs Across North America and Europe," July 27, 2026. https://www.thehackettgroup.com/the-hackett-group-finds-sga-costs-reach-five-year-highs-across-north-america-and-europe/ (I)
  2. servicedeskagents.com, "AI Service Desk Deflection Rates 2026: Benchmarks Reconciled," 2026, which traces the circulating figures to their claimed sources and finds them absent. https://servicedeskagents.com/deflection-rates/ (I, commercially motivated but verifiable method)
  3. METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," July 2025, and "Uplift update," February 24, 2026. https://arxiv.org/abs/2507.09089 ; https://metr.org/blog/2026-02-24-uplift-update/ (I)
  4. Cui, Demirer, Jaffe, Musolff, Peng and Salz, "The Effects of Generative AI on High-Skilled Work," Management Science, February 27, 2026. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4945566 (I)
  5. Faros AI, The AI Productivity Paradox 2026, 22,000 developers. https://www.faros.ai/research/ai-acceleration-whiplash (I, vendor telemetry)
  6. Google Cloud, "Announcing the 2025 DORA report," September 23, 2025. https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report (I)
  7. High Alpha, 2025 SaaS Benchmarks Report, https://www.highalpha.com/blog/is-your-team-overstaffed-for-the-ai-era-b2b-saas-company-benchmarks (I); Vista Equity Partners, AI Impact: Vista Portfolio 2026 Mid-Year Report, July 16, 2026, https://www.vistaequitypartners.com/insights/ai-impact-vista-portfolio-2026-mid-year-report/ (V, portfolio)
  8. GitClear, "The Maintainability Gap," January 2026. https://www.gitclear.com/the_ai_code_quality_maintainability_gap (I, vendor dataset)
  9. Nikolov et al., "How is Google using AI for internal code migrations?", arXiv:2501.06972, January 12, 2025. https://arxiv.org/abs/2501.06972v1 (I)
  10. AWS DevOps Blog, "Amazon Q Developer just reached a $260 million dollar milestone," August 1, 2024. https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone (V, internal estimate)
  11. Momentum, Ramp and ScyllaDB customer stories, https://www.momentum.io/customer/ramp ; Clay, A-LIGN case study, https://www.clay.com/customers/a-lign (V)
  12. Jasper, WalkMe and Akbank case studies. https://www.jasper.ai/case-studies/walkme ; https://www.jasper.ai/case-studies/akbank (V)
  13. APQC monthly-close benchmarks via Rand Group, June 1, 2026. https://www.randgroup.com/insights/services/how-long-should-month-end-close-take-benchmarks-red-flags-and-best-practices/ (I)
  14. CFO Dive, "AI cuts monthly financial close time by 7.5 days: MIT/Stanford study," August 13, 2025. https://www.cfodive.com/news/ai-cuts-monthly-financial-close-time-75-days-mit-stanford-study-accounting-accountants/757610/ (I)
  15. Numeric, "Numeric raises $51M Series B," November 19, 2025. https://www.prnewswire.com/news-releases/numeric-raises-51m-series-b-expanding-from-close-management-to-comprehensive-finance-platform-302619774.html (V)
  16. Pigment, Supercell customer story. https://www.pigment.com/customer-stories/supercell (V)
  17. MIT NANDA, The GenAI Divide, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf (I, methodology contested)
  18. Vic.ai, "Why Vic.ai"; Brex, "Accelerate accounting from transaction to close"; Ramp, "Accounting Agent launch," February 12, 2026. https://www.vic.ai/why-vic-ai ; https://www.brex.com/journal/accelerate-accounting-transaction-to-close ; https://ramp.com/blog/accounting-agent-launch (V)
  19. Ironclad, Poshmark customer story, July 14, 2023, and Hormel webinar, https://ironcladapp.com/resources/customer-stories/poshmark (V); Ironclad, 2025 Contracting Benchmark Report, https://ironcladapp.com/resources/reports/2025-contracting-benchmark-report (V, own customer base)
  20. Schwarcz, Manning, Prescott, Barry, Cleveland and Rich, "AI-Powered Lawyering," SSRN 5162111, revised May 27, 2026. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5162111 (I)
  21. CLOC, 2026 State of the Industry, March 2, 2026, 135 departments, https://cloc.org/newsdesk/cloc-releases-2026-state-of-the-industry-report-rising-legal-demand-outpaces-budget-and-staffing-growth-forcing-operational-shift/ (I); Thomson Reuters, State of the Corporate Law Department 2026, March 24, 2026, https://www.thomsonreuters.com/en-us/posts/corporates/state-of-the-corporate-law-department-report-2026/ (I)
  22. Thomson Reuters Institute, Future of Professionals Report 2026, June 22, 2026, n=1,816, via LawSites. https://www.lawnext.com/2026/06/thomson-reuters-future-of-professionals-report-warns-of-widening-gap-between-ai-adoption-and-ai-value.html (I)
  23. Vanta healthcare page (85%, uncited), https://www.vanta.com/solutions/healthcare (V); IDC, The Business Value of Vanta, January 2025, sponsored, https://www.vanta.com/resources/idc-highlights-the-business-value-of-vanta (V, sponsored); Forrester, The Total Economic Impact of Drata, October 2025, commissioned, https://tei.forrester.com/go/Drata/automationplatform (V, sponsored)
  24. Vanta, State of Trust 2025, Business Wire, October 29, 2025, 3,500 respondents. https://www.businesswire.com/news/home/20251029144534/en/Vanta-State-of-Trust-2025-AI-Threats-Outpace-Security-Expertise (V, commissioned survey)
  25. Vanta, "Vanta crosses $300M in ARR," April 29, 2026 (Samsara). https://www.vanta.com/resources/vanta-crosses-300m-in-arr-as-growth-accelerates (V)
  26. EIOPA, Digital Operational Resilience Act, applied from January 17, 2025. https://www.eiopa.europa.eu/digital-operational-resilience-act-dora_en (I, regulator)
  27. Paradox, 7-Eleven case study. https://www.paradox.ai/case-studies/7-eleven (V)
  28. Workday, "Agent Adoption Is Up 35% in One Quarter," September 1, 2026. https://blog.workday.com/en-us/leading-brands-drove-value-workday-ai-q2.html (V, named customers)
  29. Holland & Knight, "Federal Court Allows Collective Action Lawsuit Over Alleged AI Hiring Bias," on Mobley v. Workday, N.D. Cal., May 16, 2025. https://www.hklaw.com/en/insights/publications/2025/05/federal-court-allows-collective-action-lawsuit-over-alleged (I, court ruling via law firm)
  30. Moveworks, Databricks customer page, August 13, 2025. https://www.moveworks.com/us/en/customers/how-databricks-scaled-support-with-extreme-automation (V)
  31. SolarWinds, 2025 State of ITSM, October 21, 2025, 60,000+ records. https://www.businesswire.com/news/home/20251021557339/en/New-SolarWinds-Report-Gen-AI-Significantly-Drops-Incident-Response-Time (V, aggregate telemetry)
  32. Bono, Xu and Grana, "Generative AI and Security Operations Center Productivity," Microsoft, November 2024, arXiv:2411.01067; Microsoft Security Copilot randomized trials, January 2024. https://arxiv.org/abs/2411.01067 (V, self-run)
  33. ServiceNow, acquisition of Moveworks, March 10, 2025. https://newsroom.servicenow.com/press-releases/details/2025/ServiceNow-to-extend-leading-agentic-AI-to-every-employee-for-every-corner-of-the-business-with-acquisition-of-Moveworks-03-10-2025-traffic/default.aspx (I, primary release)
  34. Zylo, 2026 SaaS Management Index, January 29, 2026. https://zylo.com/news/2026-saas-management-index (I, vendor dataset)
  35. Forrester, The Total Economic Impact of Zip, Business Wire, May 26, 2026, commissioned. https://www.businesswire.com/news/home/20260526756987/en/Forrester-Study-Finds-Zips-AI-Platform-Delivers-386-ROI-for-the-Worlds-Largest-Enterprises (V, sponsored)
  36. The Hackett Group, "Rapid progress in procurement's AI agenda," March 17, 2026. https://www.thehackettgroup.com/the-hackett-group-reports-rapid-progress-in-procurements-ai-agenda/ (I)
  37. FTI Consulting, 2026 Private Equity Value Creation Index, June 4, 2026, and 2026 Private Equity AI Radar, May 19, 2026. https://www.fticonsulting.com/about/newsroom/press-releases/ai-speeds-up-returns-in-private-equity-as-ma-becomes-top-value-generator-for-firms ; https://www.fticonsulting.com/insights/reports/2026-private-equity-ai-radar (I)
  38. Gartner, June 25, 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 (I)
  39. NIST, AI Risk Management Framework 1.0, January 2023. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf (I, standard)
  40. SaaS Capital, "Spending Benchmarks for Private B2B SaaS Companies," June 10, 2026. https://www.saas-capital.com/blog-posts/spending-benchmarks-for-private-b2b-saas-companies/ (I)

Boring AI

AI for manufacturers, operators and service businesses — not startups chasing hype. Every other week.

Your address is used only to send this newsletter. No sharing, no selling, no tracking pixels. Unsubscribe from any issue.