The fixed-headcount advantage

How SMBs are growing more profitably with AI — without adding headcount

A first-party benchmark from Numa · By Asa Cox, CEO, Arcanum AI

A first-party benchmark from Numa, built on 90 days of real usage across 127 small and mid-sized businesses and the 632 people inside them who actually used it.


The businesses getting the most out of AI aren't in a luckier industry than the rest. They've learned to do a few specific things the others haven't, and the gap between the two groups is wide, consistent, and the kind any business could close.

We can say that with confidence because this isn't a survey or a vendor's best-case demo. It's a picture of how 127 real businesses actually used AI over three months: how many conversations and agents, what kinds of systems they connected, how often their teams came back.

Over one quarter, from 17 March to 15 June 2026, businesses using Numa had 27,583 conversations with it: 18,956 ordinary chats, 5,701 on-demand runs of purpose-built agents, and 2,926 agents that ran themselves, unattended, on a schedule. Behind those sat 1,663 agents these businesses built for their own jobs, 124,574 messages, and 328 business systems wired in so the AI could reach the work: the accounting system, the CRM, email, the file store.

That is the dataset: ordinary New Zealand and Australian SMBs, doing ordinary work.

The one-line version. Across 127 businesses and 632 active users, 90 days of real usage (27,583 conversations in all) models out to roughly 7,000 reclaimed human hours in a single quarter. The businesses getting the most have one thing in common, and it isn't their industry: they build AI into how the work actually runs, instead of using it as an occasional chatbot.


27,583
AI conversations (measured)
~7,091
hours saved / quarter (modelled)
~4×
more likely to lead if you connect a system
~NZ$495k
modelled value / quarter @ NZ$70/hr

Two kinds of number, and we never blend them

One rule sits under every figure below, because it's the rule that makes them worth trusting.

Measured numbers are hard counts: conversations, agents, runs, active users, systems connected, cost. A sceptic can screenshot any one of them and we'll stand behind it on its own.

Modelled numbers are estimates. Chief among them is hours saved. For most work there's no stopwatch: nothing records that a task 'would have taken Jane 40 minutes by hand.' Some businesses do put their own time-saved estimate on the agents they build, and we use that as a cross-check, but it only covers a slice of the work. So the fleet-wide figure is modelled, transparently, and always shown as a range rather than a single confident number.

We never present a modelled hour as a measured fact, and we never quietly merge the two. Where you see 'roughly 7,000 hours,' it's modelled, it carries a range, and we tell you exactly how we got it (see the box at the end). The activity counts underneath it are measured and exact. That discipline matters, because the findings that matter most here are the measured ones.


The spine: roughly 7,000 hours back in a quarter

Translate those 27,583 conversations into human time. We multiplied each one by how long the equivalent work would plausibly have taken a person by hand, an estimate derived from a corpus of 12,464 classified conversations, banded low / central / high, with a haircut applied because not every completed task genuinely saved someone time.

The central estimate: about 7,091 hours of human time saved across the fleet in 90 days (modelled; range roughly 2,200–20,600; n = 127 businesses).

The range is wide on purpose; a single confident number would be an inaccurate version. Even at the most conservative end, it's thousands of hours a quarter that nobody had to work. There's an independent cross-check, too: a separate, more conservative tally built only from the time-savings businesses themselves set on their scheduled agents comes to 922 hours for the scheduled slice alone, which sits comfortably inside our scheduled-agent band.

Hours saved is only the starting point. What those hours free up is the part that matters: a quote out the door before a competitor's, a month-end that doesn't run into the weekend, an extra job won without another hire. The gain is capacity more than cost savings: a fixed team getting more done.


Where the reclaimed hours come from
Modelled hours saved over 90 days, by interaction type
Plain chat4,142On-demand agent2,082Scheduled agent868

The size of the prize: how far the leaders have pulled ahead

The most useful pattern in the data is also entirely measured.

Among the businesses that actively used Numa over the quarter, the typical one ran around 110 conversations, and the most engaged ran into the thousands. The leaders have made Numa part of how the work gets done, and they pull far more out of it than a business that has only tried it a few times.

That distance is the opportunity. A small group of businesses is capturing most of the available value today, which means most of it is still available to everyone else. The rest of this piece is about what those leaders do, since every move is one any business can make.

And the ranking isn't an artefact of the model. Order the businesses by modelled hours and you get almost the same order as ranking them by raw conversation count. The leaders are simply the ones doing the most real work with Numa.


What the leaders do differently, and it isn't their industry

So what separates the businesses pulling ahead? We profiled the top quartile against everyone else. The answer is three behaviours, all measured, all copyable, plus one popular theory that turns out to be wrong.

1. They run their work through agents. The leaders push 47% of their AI work through purpose-built agents (named tools for recurring jobs) versus 15% for everyone else.

2. They wire in their own systems. Top-quartile businesses connect a median of 6 of their own systems (Xero, the CRM, email, the file store) versus 3 for the rest.

3. They use it harder and build more. Each active user in a leading business runs about 35 conversations a month versus 7 for the rest, and they've built a median of 19.5 agents versus 12.

The theory that's wrong is the one about breadth. Both groups already use AI across roughly a dozen different job types. The leaders go deeper and more automated on the jobs they already do. The move is to take the use-cases you already have and turn them into connected, automated agents.

The biggest 'industry effect' in the data largely disappears once you account for these behaviours. Company size has essentially zero relationship with per-person value (rank correlation −0.09). The real predictors are usage intensity, systems connected, and whether the business has automated anything.

'The industry doesn't predict the payoff. The workflow does. How deep you wire AI in, and how hard your team uses it, is what separates the leaders from everyone else.'

This is the answer to the first question every SMB owner asks: 'Will this even work for my kind of business?' The honest answer from the data is that it's the wrong question. It works for the businesses that use it like the leaders do, in every sector we can report.

What the leaders do differently — and it isn't their industry
Top quartile vs everyone else, across four measured behaviours
Leaders (top quartile)Everyone else
% of AI work via agents47%15%Own systems connected (median)63Conversations / user / month357Agents built (median)19.512

The one move that predicts a winner

We went back and asked which early behaviour actually predicts a business ending up in the top quartile, and one signal beat everything else.

It isn't building an agent. Almost every active business does that: 97% of new customers build at least one agent, most in the first week. Because everyone does it, it tells you nothing about who's going to win.

What predicts a winner is connecting a system. A business that wires in even one of its own tools is about 4× more likely to land in the top quartile than one that never connects anything (33.9% versus 8.0%). It's the strongest leading indicator in the whole dataset, and it holds up both across the full base and in a cohort of genuinely new customers tracked from day one, where it still lifts the odds about 3.3×. The cliff is stark: 90.5% of top-quartile businesses have connected a system, against just 6.5% of the least active.

The runner-up is scheduling. More than half of the businesses that turn on a single automated schedule end up in the top quartile.

That gives customer success an honest scoreboard. The question isn't 'did they build an agent.' Everyone builds an agent. The question is 'have they connected Xero or their email or their drive yet, and is anything running on a schedule.' A business 30 days in with zero systems connected is the one to call first.


The step-change at the third system

If the benchmark leaves an owner with one thing to go and do, it's this.

Value doesn't climb smoothly as you connect each integration; it jumps. One or two connected systems make almost no difference. At the third, hours saved per person nearly quadruples.

The connected workflow is what does it. With one app wired in, AI is still a smart assistant sitting alongside the business. By the third system it's working inside it: pulling the job data, cross-referencing the accounting system, and writing the result back without anyone ferrying files between tabs.

Connected businesses save roughly 4.5× more per person than unconnected ones. A sceptic will say that's just because connected businesses are more engaged anyway. Fair point, and we tested it. Even holding engagement constant, the integration-only effect is still about 1.7×.

There's a natural sequence that doubles as a playbook: connect your systems first, then automate on top of them. Across the dataset, almost nobody scheduled an agent before connecting systems.


Value jumps at the third connected system
Median hours saved per active user per month, by systems connected (modelled)
None0.731–2 systems0.743–5 systems2.806+ systems3.97

What scheduled agents add

A scheduled agent runs on its own, every morning say, with no one watching. Across the fleet, a single schedule ran a median of 26 times in 90 days (typically 12 to 49, up to 129). The setup happens once, and the agent does the job all quarter.

Businesses running schedules save about 3× the hours per person of those that aren't, though schedulers tend to be the more advanced accounts already, so scheduling goes with depth of use rather than causing it on its own.

What people schedule is the genuinely repeatable work. 95% of scheduled-agent activity is automation and operations (ticket triage, the morning inbox-and-calendar sweep, a recurring financial projection), with only 0.7% throwaway, against nearly a third for plain chat.


The same three jobs, everywhere

When you ask what AI is actually doing for these businesses, the answer is reassuringly consistent across every sector we can report. Three jobs drive about 59% of all the hours saved:

The overall split is close to even across the three: automation and operations 37%, documents and content 32%, analysis and technical 30%.

Industry changes how much AI gets used far more than what it's used for. Construction leads on building automations. Professional services leads on analysis and adds tender and compliance work as a fourth area.


See these in practice. The behaviours in this benchmark map to real agents SMBs run on Numa: estimating & quote generation, job & project profitability, meeting analysis & action items, internal knowledge assistant, accounts payable automation and financial analysis & forecasting — or browse all 50+ in the Numa use-case library.

Where AI isn't delivering yet

A benchmark that only reported its wins wouldn't be worth much, so here is where the data is less flattering.

The least valuable work is also the most common. Over two-thirds of all activity (69%) is one-shot plain chat, while scheduled automation is only about 11%. The biggest missed opportunity is all the one-off chats that could have been turned into reusable agents.

There's also one thing we couldn't measure: who, by role, saves the most. Rather than guess at it, we've left it out, and it's the main reason the next benchmark needs a short customer survey.


What building with AI actually looks like

The builder who quotes faster. In construction and trades, the most common pattern is turning site paperwork into priced work. Businesses build agents that take the PDFs and drawings, pull the quantities into an estimating take-off, and drop them into the company's own spreadsheet. Others rebuild a job-profitability report to split materials from labour. It's the textbook capacity story: the same crew, the same estimator, more jobs quoted and out the door.

The professional-services firm that automates the morning. A common build here is a scheduled, unattended briefing agent. It triages the inbox overnight and produces a morning summary through the firm's connected email, or distils a board pack down to the few points that matter before a meeting. Multiply that by 26 unattended runs a quarter and the leaders' numbers start to make sense.

The common thread is a change in how these businesses use AI: from asking it one-off questions to having it build the tools that do the work on their own.


What it's worth (read this as illustrative)

We've deliberately kept dollars out of the headlines, because they rest on assumptions a sceptic can wave away. But the value is real. At an illustrative loaded labour rate of NZ$45–110/hour (central NZ$70, salary plus on-costs), the reclaimed hours across all 127 businesses model out to roughly NZ$495,000 of value in the quarter (band ~NZ$100k–$2.3M; central shown), on the order of NZ$2.0M a year. Figures are in New Zealand dollars; Australian dollars sit within about 10%.

Treat the dollar figure as directional: it's modelled, and it shifts with whatever loaded labour rate you consider realistic. The hours underneath it are the firmer number.


The takeaway for an owner

  1. Industry isn't the deciding factor. The high-value businesses are spread across trades, manufacturing and professional services. What sets them apart is how they use AI.
  2. Use it for the work itself, not only for quick answers. The leaders run real, recurring jobs through Numa rather than treating it as a chatbot.
  3. Connect your systems, and get past the third one. It's the behaviour that most reliably predicts a high-value account.
  4. Build agents, then put them on a schedule. The most valuable hour you'll spend is the one that builds something that then runs on its own every day after.

How we measured this

Counts (conversations, agents, businesses, active users, systems connected, runs, cost) are measured hard counts over a fixed 90-day window (17 March to 15 June 2026). Hours saved are modelled: each conversation's measured type is multiplied by a corpus-derived minutes-per-conversation estimate (low / central / high), plus a 'realisation haircut'. We report the central estimate with the low–high range always attached, and never present a modelled hour as measured. Dollar figures are modelled on modelled — modelled hours times an illustrative NZ$45–110/hr rate (central NZ$70) — and are deliberately illustrative. Headline distributions use the median. Leading-indicator claims are predictive, not causal. Every reported segment has at least 5 businesses. The benchmark population is 127 curated external customers. Every hero figure was independently re-derived from source data with zero discrepancies.


Frequently asked questions

Does a business's industry determine whether AI pays off?
No. In a first-party benchmark of 127 SMBs over 90 days, industry barely predicted the payoff from AI. The gap between high-value and low-value businesses was driven by copyable behaviours — how deeply AI was wired into their systems and how hard their teams used it — not by sector.

What most predicts whether an SMB gets value from AI?
Connecting one of your own systems. A business that wires in even one tool (accounting, CRM, email, file store) was about 4× more likely to become a top-value account than one that connected nothing (33.9% vs 8.0%) — the strongest leading indicator in the dataset.

How much time can AI save a small or medium business?
Across 127 businesses over one 90-day quarter, 27,583 measured AI conversations modelled out to roughly 7,091 reclaimed human hours (modelled central estimate; range ~2,200–20,600). Conversation counts are measured; hours saved are transparently modelled and shown as a range.

What do the businesses getting the most from AI do differently?
Three measured behaviours. They run work through purpose-built agents (47% of their AI work vs 15% for the rest), they connect more of their own systems (a median of 6 vs 3), and they use it harder (about 35 conversations per active user per month vs 7).

Why does connecting a third system matter so much for AI value?
Value from AI doesn't climb smoothly — it jumps at the third system. Median modelled hours saved per active user per month go from 0.73 (none) and 0.74 (1–2 systems) to 2.80 at 3–5 systems and 3.97 at 6+.

What is Numa?
Numa by Arcanum AI is an AI operating system for SMBs. It connects the tools a business already runs — Microsoft 365, Xero, MYOB and 2,000+ apps — and automates the admin across them with AI agents built in plain language. It is built by Arcanum AI, a New Zealand company founded in 2016, and is distinct from the unrelated Numa serving car dealerships (numa.com).