Chapter 3
Gross margin: the six levers below the revenue line
Support, implementation, services, hosting and the cost of the AI itself. This is where the margin structure of a software company changes, and where the best-evidenced savings in the whole book sit next to the one cost line that AI adds rather than removes.
Chapter 3 of 10Gross margin13 min
Gross margin is where Deploy becomes Reshape. The tools that resolve tickets, migrate data and rightsize infrastructure are purchases; the reorganization of the support, implementation and services functions behind them is the work. Two things distinguish this chapter from the operating-expense chapter that follows. The savings here are larger per lever, because cost of goods scales with customers rather than with the company. And one lever, G5, runs the other way: inference is a new cost of goods, and a company that ships AI product without managing it can grow revenue and shrink gross profit in the same year.
G1 · Support resolution
What it is. A tier-one agent grounded in the knowledge base, product documentation and the customer's own configuration and usage data, resolving how-to, access, configuration and report questions in chat and email, measured on true resolution (the customer confirms or does not return within seven days) rather than containment (no human handoff), with an assist agent on every ticket a person handles, a knowledge agent maintaining the articles the resolution agent runs on, and a telemetry agent opening tickets before customers notice. Voice follows once chat and email are stable.
The line it moves. Support and success cost of goods; first-response time and CSAT as retention inputs.
Typical range and the evidence. Strong. Support cost from around 10% of revenue to around 5% over three years, with resolution by agent reaching 50% to 60% at maturity and headcount falling 40% to 50% through attrition, is the case-study range. Intercom reports its Fin agent "averaging 76% across 12,000+ customers," and the definition matters: a soft resolution counts when the customer "exits the conversation without requesting further assistance within 24 hours," so part of that figure is silence.1 Lorikeet's 2026 synthesis of vendor data says to "expect 40% to 60% at launch and 60% or more after 6 to 12 months of knowledge and workflow work," and notes that "containment runs about 20 points above resolution on the same conversations."2 HubSpot's customer agent "resolves 72% of support tickets without human escalation" at a self-serve scale.3 Klarna is the caution: in February 2024 its assistant "had 2.3 million conversations, two-thirds of Klarna's customer service chats," doing "the equivalent work of 700 full-time agents," and by May 2025 its chief executive was telling Bloomberg that "what you end up having is lower quality" and that "really investing in the quality of the human support is the way of the future for us."4 The independent evidence on augmentation is stronger than on replacement: Brynjolfsson, Li and Raymond's study of 5,179 agents found a generative assistant "increases productivity, as measured by issues resolved per hour, by 14% on average, including a 34% improvement for novice and low-skilled workers."5
Conditions. A single ticketing system with clean history; the knowledge base rebuilt in the first sixty days; read access to configuration and usage data; an evaluation set of a few hundred real tickets with known-good answers that the agent passes at 85% before launch and after every model or prompt change; a covered model endpoint where tickets contain regulated data; a human escalation path staffed at all hours the product is used. A purchase, not a build.
Kill criteria. True resolution below 40% six months after launch; CSAT below baseline for two consecutive months; any month in which the fifty lowest-scored conversations show the same failure twice.
- Tickets resolved by agent (true resolution)Support and successAt close0%Year 140%Year 360%
- Cost per ticketSupport and successAt closeBaselineYear 1−35%Year 3−60%
- First response timeRetentionAt closeHoursYear 1MinutesYear 3Under 2 minutes
- CSATRetentionAt closeBaselineYear 1Baseline or betterYear 3Baseline +5 pts
- Support and success, % of revenueGross marginAt close8 to 10%Year 16 to 7.5%Year 34 to 5%
Horizon and stage. Deploy, then Reshape when the team is resized and merged with success. Day one hundred through year three.
How it translates. A: phone remains the channel; the voice agent is the year-two step. B: tickets are regulatory questions, and the knowledge base is the product's own content. C: high volume from small merchants; the largest absolute saving of the five. D: the HubSpot pattern, at 72%. E: low volume, high complexity; the assist agent matters more than the resolution agent.
G2 · Onboarding and implementation
What it is. Data migration by agent (reading exports, mapping fields, flagging conflicts, producing a validated import with human review of exceptions only), configuration by interview, guided self-serve onboarding for smaller customers, training generated from the customer's own configuration, and a go-live readiness model that directs the remaining implementers to the accounts that need them. Executed after the pricing decision that bundles implementation into the subscription.
The line it moves. Implementation cost of goods; time to value, which is the leading cause of early churn; revenue recognition timing.
Typical range and the evidence. Medium, and the thinnest evidence base among the well-known levers. Time to go-live halved and implementation hours per customer down 40% to 60% is the case-study range, run as an internal measurement program. BuildOps, a field-service vertical company, used a machine-learning import tool in 2022 to "decrease its overall time to launch by an estimated 15%-20%" and record an "average 20%-25% decrease in per-project person hours."6 Workday's Deployment Agent is "designed to deliver an estimated 30% reduction in implementation hours and costs" on current projects, targeting "up to 50%," with named customers and the hedge in the vendor's own words; in the same quarter Workday guided professional services revenue down, which is what this lever looks like from the top line.7 Guidewire says its "AI-powered project harness for implementations is delivering on the promise of material reduction in project complexity and duration," with no number attached.8 No vendor publishes a clean before-and-after on implementation cost per customer.
Conditions. A documented target schema and import API; a library of anonymized migration cases for testing; the pricing decision made first; a purchased data-import platform for the file handling, with the mapping logic built in-house because the schema is the company's own.
Kill criteria. Implementation hours per customer not down 25% after the agents have run on twenty implementations; or any migrated record set that fails count-and-sample reconciliation before go-live.
- Time to go-live (median)Revenue timing, early churnAt close8 weeksYear 14 weeksYear 32 weeks
- Implementation hours per customerProfessional servicesAt closeBaselineYear 1−40%Year 3−60%
- Customers onboarded per implementer per monthProfessional servicesAt close2Year 14Year 37
- Implementation cost recoveryRevenue, PS COGSAt closeBelow cost, billed hourlyYear 1Bundled in tierYear 3Bundled in tier
Horizon and stage. Reshape. Years one to two.
How it translates. A: client records, caregiver credentials, authorizations and payer setups migrated under a business associate agreement. B: the customer's product data and prior determinations loaded, which is where the content moat is tested. C: merchant setup and payments onboarding, where speed to first transaction is the metric. D: onboarding is the product (R11). E: the lever that carries the bridge, with G3.
G3 · Professional services margin
What it is. Services repriced or productized: implementation into the subscription, configuration into product (R2), training into generated content, and the residual services line run at a margin rather than as a loss leader. For a services-heavy business it is the whole plan.
The line it moves. Professional services gross margin; revenue mix; and the multiple, because a dollar of services revenue is valued at a fraction of a dollar of subscription revenue.
Typical range and the evidence. Strong, because the gap is in public segment disclosures. Guidewire's fourth quarter of fiscal 2026: "Services' gross margin was 12.5% compared with 12.9% a year ago," against "Subscription and support gross margin was 74.5%, up 4 percentage points year-over-year," on the same customers; services revenue grew 23% while the margin slipped, which is the investment preceding the return.8 Benchmarkit's 2025 private benchmarks put professional services at about 15% of revenue and about 30% margin, and KeyBanc's older survey at about 26%.9 The Workday guide-down above is the mirror image: successful implementation AI shows up as shrinking services revenue, which looks like weakness in the top line and is mix improvement.
Conditions. A services P&L that is measured (most are not); a pricing decision that moves implementation into tiers for new customers and, at renewal, for existing ones; a partner strategy where the services are large enough to hand to integrators on fixed-price terms, which Guidewire says its partners will offer "as they get more confident in their tooling."8
Kill criteria. Services revenue falling faster than subscription rises for two consecutive quarters, which means the bundling was priced too low.
- Services, % of revenueRevenue mixAt close15 to 35%Year 1−5 pointsYear 3Below 20%
- Services gross marginGross marginAt close10 to 30%Year 1+5 pointsYear 335 to 45%
- Implementation bundled into subscription (new customers)RevenueAt close0%Year 1100%Year 3100%
- Subscription revenue growth net of services conversionRevenueAt closeBaselineYear 1RisingYear 3Rising
Horizon and stage. Reshape. Years one to three.
How it translates. A: a small services line billed below cost; bundled in year one. B: services are the content onboarding, small. C: negligible. D: none. E: the bridge itself, from 35% of revenue at 15% margin toward 20% at 40%.
G4 · Hosting and FinOps
What it is. Continuous rightsizing, reserved-capacity purchasing, idle-resource cleanup and workload scheduling run by tooling under the FinOps Foundation's framework, "an operational framework and cultural practice which maximizes the business value of technology ... through collaboration between engineering, finance, and business teams."10
The line it moves. Hosting cost of goods.
Typical range and the evidence. Strong. Fifteen to twenty-five percent of the hosting bill is what to plan on. Flexera's 2026 survey of more than 750 cloud decision-makers found wasted infrastructure and platform spend rose to 29%, the first increase in five years, driven by AI workloads.11 Cast AI's Akamai case quotes savings "falling between 40-70%, depending on the workload," a per-workload figure rather than a total-bill one.12 The lever is a purchase and it pays inside a quarter.
Conditions. Tagging and cost allocation across cloud accounts; observability; documented runbooks; an approval step on any production change for the first six months.
Kill criteria. Hosting cost per dollar of revenue not down 10% after two quarters.
- Hosting cost per $1 of revenueHostingAt closeBaselineYear 1−6%Year 3−10 to −15%
- Share of spend under commitment or reservationHostingAt closeLowYear 150%Year 370%
- Idle or untagged resourcesHostingAt closeUnknownYear 1Under 5%Year 3Under 2%
Horizon and stage. Deploy. Day one hundred to year one.
How it translates. Roughly the same everywhere; largest in absolute terms for D, where hosting is a large line, and for E, where multiple acquired stacks duplicate infrastructure until X4 consolidates them.
G5 · Inference cost management
What it is. Model routing by task, small-model and cached fallbacks, multi-provider architecture, retrieval that reduces tokens per answer, and per-feature cost of goods tracked through the gateway (P7), so that the AI product (R2) and the internal agents do not consume the margin they create.
The line it moves. Hosting cost of goods, and the gross margin of the AI product specifically; and the exit story, because buyers grade software on gross margin.
Typical range and the evidence. Strong for the price trend, medium for the practice, and the two pull in opposite directions. Epoch AI found the price of a fixed capability level falling "ranging from 9x to 900x per year," with GPT-4-level performance on graduate-level science questions falling from $37.50 per million tokens in March 2023 to $0.12 in December 2024, where the data ends.13 That measures yesterday's capability; frontier capability has not fallen the same way, and agentic workloads consume far more tokens per task. Growth Unhinged's 2026 survey puts the median target gross margin for AI products at about 50% against 70% to 80% for traditional software, six years after a16z put AI companies at "gross margins often in the 50-60% range."14 The practice evidence is that the cost is controllable: Atlassian's agents on its knowledge graph deliver "up to 44% more accurate answers while consuming 48% fewer tokens," a gross-margin disclosure disguised as a product metric;15 Notion's chief financial officer reported cutting AI infrastructure cost about threefold over two years through multi-model routing;16 and Figma reported non-GAAP gross margin of 85%, up 2.5 points sequentially, while shipping AI aggressively, naming task-based routing, a model-agnostic architecture and first-party models as the three levers, with the caution that margin "will vary quarter-to-quarter in the near term" as beta products consume inference without revenue.17 The a16z partner's view that 85% to 90% gross margin is "an orange flag" because "there's probably not much AI usage in your product" is the contrarian position, and for a seller both views are true at once.16
Conditions. Every model call through the gateway with cost attributed to a feature or agent; a routing policy that defaults to the cheapest model that passes the evaluation set; AI product priced with its cost of goods known.
Kill criteria. Any AI feature whose cost of goods exceeds 50% of its attributable revenue for two quarters is repriced or re-routed; if neither works, it is withdrawn.
- Inference cost, % of revenueHostingAt close0 to 2%Year 1Under 4%Year 3Under 5%
- AI product gross marginGross marginAt closen/aYear 150%Year 365 to 75%
- Share of model calls on non-frontier modelsHostingAt closeUnknownYear 160%Year 380%
- Blended gross marginGross marginAt closeBaselineYear 1HeldYear 3Held or better
Horizon and stage. Deploy. Years one to three, continuously.
How it translates. A: small in dollars; matters for the documentation module. B: the rules engine should be deterministic where it can be, with the model on the edges, which is both cheaper and auditable. C: modest. D: the lever that decides whether the bridge holds, because the AI product is the growth story and inference is its cost. E: modest.
G6 · Cost to serve
What it is. Customer success and account management restructured around the signals and agents from R4 and R5: the renewal brief assembled by agent ninety days out, the quarterly review generated from live data, the voice-of-customer digest, and a merged customer team where support and success were separate.
The line it moves. Support and success cost of goods, as distinct from the support-ticket cost G1 removes.
Typical range and the evidence. Medium. ARR per success manager up 15% to 25% over three years is the case-study range, and Vista's seventeen-company figure (from $5.4M to $6.7M) is the anchor.18 LogicMonitor's "60%+ reduction in response time" is a Vista portfolio figure from the same report.18 What no vendor publishes is the cost-to-serve delta from merging the functions, so the merge is run as an internal measurement.
Conditions. R4 and R5 in place; a single customer team leader; compensation on net retention; the pricing council's rules on renewal recommendations.
Kill criteria. Gross retention falling in the quarter after the merge, in which case the functions are separated again and the merge is retried a year later.
- ARR per success managerSupport and successAt closeBaselineYear 1+5%Year 3+15 to 25%
- Hours to prepare a renewal briefSupport and successAt closeDaysYear 1HoursYear 3Minutes plus review
- Support and success combined, % of revenueGross marginAt close8 to 12%Year 1−2 pointsYear 3−4 to −5 points
Horizon and stage. Reshape. Years one to two.
How it translates. A: support and success become one customer team of nine, from twelve. B: account management is regulatory advisory; the digest matters most. C: high-volume, low-touch; the lever is automation of the renewal itself. D: self-serve; minimal. E: named account teams; the brief and the review are the gains.
Sources
- Intercom, fin.ai (accessed September 2026); "From resolutions to outcomes," March 12, 2026; resolution definitions from Intercom staff on the Intercom community forum. https://fin.ai/ ; https://www.intercom.com/blog/from-resolutions-to-outcomes-evolving-how-fin-delivers-value/ (V, anonymous)↩
- Lorikeet, "AI customer support resolution rate benchmarks 2026," updated September 2, 2026. https://www.lorikeetcx.ai/articles/resolution-rate-ai-customer-support-benchmarks-2026 (V, synthesis)↩
- HubSpot Q2 2026 earnings call, August 12, 2026. https://www.fool.com/earnings/call-transcripts/2026/08/12/hubspot-hubs-q2-2026-earnings-call-transcript/ (I, public company)↩
- Klarna, PR Newswire, February 27, 2024, https://www.prnewswire.com/news-releases/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month-302072740.html (V); Sebastian Siemiatkowski to Bloomberg, May 8, 2025, via Fortune and Customer Experience Dive, https://www.fortune.com/2025/05/09/klarna-ai-humans-return-on-investment (I)↩
- Brynjolfsson, Li and Raymond, "Generative AI at Work," Quarterly Journal of Economics 140(2), 2025. https://www.nber.org/papers/w31161 (I)↩
- Flatfile, "BuildOps clears data import bottleneck with Flatfile," 2022 implementation. https://flatfile.com/resources/case-studies/buildops-clears-data-import-bottleneck-with-flatfile/ (V)↩
- Workday Q1 fiscal 2027 earnings call, May 21, 2026. https://www.fool.com/earnings/call-transcripts/2026/05/21/workday-wday-q1-2027-earnings-call-transcript/ (I, public company; design target, not audited result)↩
- Guidewire Q4 fiscal 2026 earnings call, September 9, 2026. https://www.fool.com/earnings/call-transcripts/2026/09/09/guidewire-gwre-q4-2026-earnings-call-transcript/ (I, public company)↩
- Benchmarkit, 2025 B2B SaaS Performance Metrics Benchmarks, https://www.benchmarkit.ai/2025benchmarks ; KeyBanc 2017 private SaaS survey via forentrepreneurs.com (I)↩
- FinOps Foundation, "What is FinOps?", updated March 2026. https://www.finops.org/introduction/what-is-finops/ (I)↩
- Flexera, 2026 State of the Cloud Report, March 18, 2026, 750+ respondents. https://www.flexera.com/about-us/press-center/flexera-finds-cloud-value-is-rising-while-ai-waste-grows (I)↩
- Cast AI, Akamai case study. https://cast.ai/case-studies/akamai/ (V)↩
- Epoch AI, "LLM inference prices have fallen rapidly but unequally across tasks," March 12, 2025; data through December 2024. https://epoch.ai/data-insights/llm-inference-price-trends (I)↩
- Growth Unhinged, May 13, 2026, as in the revenue chapter (I); a16z, "The New Business of AI," February 16, 2020, updated April 25, 2024, https://a16z.com/the-new-business-of-ai-and-how-its-different-from-traditional-software/ (I, dated)↩
- Atlassian, Q4 FY26 shareholder letter, August 6, 2026. https://www.sec.gov/Archives/edgar/data/1650372/000165037226000031/teamq42026shareholderlet.htm (I, public company)↩
- Mostly Metrics (CJ Gustafson), "Can bad gross margins ever be a good sign?", November 9, 2025, reporting Sarah Wang of a16z and Notion's CFO. https://www.mostlymetrics.com/p/can-bad-gross-margins-ever-be-a-good-sign (I, named sources via independent outlet)↩
- Figma Q2 2026 earnings call, August 12, 2026. https://www.fool.com/earnings/call-transcripts/2026/08/12/figma-fig-q2-2026-earnings-call-transcript/ (I, public company)↩
- Vista Equity Partners, AI Impact: Vista Portfolio 2026 Mid-Year Report, July 16, 2026. https://www.vistaequitypartners.com/insights/ai-impact-vista-portfolio-2026-mid-year-report/ (V, portfolio)↩
Boring AI
AI for manufacturers, operators and service businesses — not startups chasing hype. Every other week.
Your address is used only to send this newsletter. No sharing, no selling, no tracking pixels. Unsubscribe from any issue.