AI for pricing decisions, margin, and business insight

AI Business Analytics for Service Business Owners: Reading Your Own Numbers Without Guessing

By Ricky West · Founder, Turnkey Services · August 7, 2026 · 15 min read

A plumbing owner I know sat down last spring convinced his drain cleaning work was carrying the company. It was the biggest line on his revenue report by a wide margin. Then we pulled fourteen months of jobs out of his field service software and spent an afternoon working through them. AI business analytics for service business owners is not a dashboard product — it is the act of pointing a language model at data you already own and asking it questions your revenue report cannot answer. Three hours later, drain cleaning was still the biggest revenue line and the fourth-most profitable job type he ran. Water heater replacements, which he thought of as occasional, were quietly producing more gross profit per truck-hour than anything else on the board.

Nothing about that afternoon was sophisticated. We did not buy anything, connect anything, or write code. We exported a CSV, cleaned it up, and asked better questions than his software was designed to answer. That is the whole method, and this piece walks through it the way it actually goes — including the parts where it fails.

Start by admitting which numbers you are currently guessing at

Every owner in the $1M to $5M range has a mental model of their business, and most of that model is built from memory and feel rather than arithmetic. That is not a criticism. You are running trucks. But it is worth writing down, on paper, the questions you answer by instinct:

Those five questions are answerable from data you already have. None of them are answered by the default reports in Jobber, Housecall Pro, or ServiceTitan, because those reports are built to summarize activity rather than interrogate profitability. Revenue by job type is activity. Gross profit per truck-hour by job type is a decision.

Pick two questions. Not ten. The owners who get value out of this start narrow and finish, and the ones who try to build a full analytics practice in one sitting abandon it by Thursday. If you are still deciding whether any of this is worth your Saturday, the honest accounting in this ROI breakdown for service owners is a reasonable place to calibrate expectations first.

The export is the hard part, and nobody warns you

Here is where the first real caveat lives. Your field service software will happily give you a jobs export. What it gives you is usually job date, customer, job type or tag, assigned technician, invoice total, and status. What it usually does not give you is labor cost per job, and without labor cost you cannot compute margin — you can only compute revenue, which you already knew.

So the first thing to check is whether your techs are actually clocking in and out per visit inside the software, or whether time is tracked in a separate payroll system that never touches the job record. In shops running three to eight trucks, it is roughly a coin flip. If your time data lives somewhere else, you have two options: export both and match them by date and tech, or use a loaded hourly rate per tech as a stand-in. The stand-in is less accurate and completely adequate for a first pass.

To build that stand-in, take a tech's base wage and multiply it by a loading factor that accounts for payroll taxes, workers' comp, general liability, phone, uniforms, and vehicle allowance. For most residential trades shops that factor lands somewhere between 1.45 and 1.8. Pick one number, apply it consistently across every tech, and write down which factor you used so your next analysis is comparable. Do not agonize over whether it should be 1.55 or 1.6 — the ranking of your job types will almost never flip on that margin, and the ranking is what you are after.

The structural difference between platforms matters too. ServiceTitan reports at the invoice-line level, so a single job can produce six rows. Jobber and Housecall Pro report at the job level, so that same job is one row with a total. If you have ever wondered why two people in your shop quote different average tickets, that is often why. Before you ask AI anything, know which grain your export is at. If you are still deciding how your tools should talk to each other, the piece on which AI tools actually connect to Jobber, Housecall Pro, and ServiceTitan covers where the real integration seams are.

What to strip before it touches a chat window

Delete customer names, street addresses, phone numbers, and email addresses from the export. You do not need them to answer margin questions, and there is no reason to put customer contact information into a general-purpose AI tool. Keep a customer ID column if you want repeat-purchase analysis — a number is enough to group by. Keep tech first names or initials only. This takes ninety seconds in a spreadsheet and it is not optional.

Clean up your job type column before anything else

This is the step that quietly determines whether the whole exercise works. Most shops that have been running software for a few years have a job type list that grew by accretion: someone added "WH Replace," someone else added "Water Heater - Replacement," a third person tagged nine of them "Install." The model will treat those as three separate categories and report three separate margins, none of which mean anything.

Open the export, sort by job type, and look at the distinct values. If you have more than fifteen or twenty for a residential trades shop, you have a naming problem, not a complexity problem. Consolidate into a short list you would actually use to make decisions — diagnostic, repair, replacement, install, maintenance, warranty — and add a second column for the system or trade if you need it. Do this once, save the mapping, and reuse it every quarter. Half the shops I have watched attempt this analysis got a nonsense answer the first time and blamed the AI, when the real cause was thirty-one job types describing eleven kinds of work.

Ask the question the way you would ask a sharp new office manager

The mistake almost everyone makes on their first attempt is asking something vague like "analyze my job data." You get back a competent, useless summary: total revenue, job counts, a few averages, and a closing paragraph about opportunities for growth. That is not analysis. That is a table of contents.

What works is giving the model the same context you would give a competent human who just walked in the door. Tell it what the columns mean. Tell it how your labor cost column was built. Tell it what you suspect. Then ask a specific question with a specific output shape.

A prompt that actually produces something useful looks closer to this:

"This is fourteen months of completed jobs from a residential plumbing company running four trucks in a metro market. Columns are job date, job type, assigned tech, on-site hours, invoice total, parts cost, and loaded_labor_cost — that last column is already fully burdened, so use it as-is and do not apply any additional multiplier. Calculate gross profit per on-site hour for each job type, rank them, and show me job count and total revenue alongside so I can see volume. Flag any job type where the median and the mean are more than 25 percent apart, because that tells me the average is hiding something. State how many rows you processed. Do not round to whole dollars."

That instruction about median versus mean is the one that earns its keep. In service work, averages lie constantly. One five-figure sewer line replacement will drag the average for an entire job category upward and convince you a thin category is healthy. When the median lands well below the mean, you are looking at a category that survives on outliers — which is a completely different business decision than a category that is consistently good.

The broader habit here — writing prompts with defined inputs, defined outputs, and defined constraints — is the same skill that makes every other AI task in your shop work. The rundown of what general-purpose AI tools actually do well for owners covers the pattern in more detail if this is new territory.

The four analyses that pay for the afternoon

Across the shops I have watched do this, four questions produce nearly all of the value. Run these before you get creative.

1. Gross profit per truck-hour, by job type

This is the reversal test. Rank your job types by revenue, then rank them by gross profit per on-site hour, and put the two lists side by side. When the order changes — and it usually does — you have found something. The classic pattern in HVAC and plumbing shops is that the highest-revenue job type falls to third or fourth once drive time, callback rate, and real parts cost get assigned to it. High-revenue, high-hour, thin-margin work feels like the backbone of the company because it is always on the schedule. That is exactly what makes it hard to see.

2. Callback and warranty work, isolated

This is the leak almost nobody measures, because callbacks are usually logged as a zero-dollar invoice attached to the original customer. Zero revenue means they vanish from every revenue report you run, while still consuming a full truck-hour, a fuel cost, and a slot that a paying job could have used. Ask the model to find every job with an invoice total at or near zero, group them by original job type and by tech, and express the result as a percentage of that job type's total on-site hours. A shop running 2,000 jobs a year with a 6 percent callback rate on installs is giving away real capacity, and it will not show up anywhere until you look for it deliberately.

One caution on the tech breakdown: your best technician will often show an elevated callback count, because dispatch sends the hard jobs and the second opinions to the person who can handle them. Look at callbacks against the difficulty mix that tech was assigned, not in isolation. This is a case where the number starts a conversation rather than settling one.

3. Quote-to-close, split by estimator and ticket band

Aggregate close rate is a vanity number. What you want is close rate split by who wrote the quote and by ticket size band. Define four bands using thresholds that mean something in your shop — a small diagnostic-and-repair band, a mid band, a band where the customer usually asks a spouse, and a top band where you personally get involved. Estimators who are excellent at small work often fall off a cliff on large work, and the aggregate number hides it completely. This analysis is also the one most likely to change how you route estimates tomorrow morning. If quoting speed is part of your problem, how contractors use AI to quote jobs faster without losing accuracy works through that side of it.

4. Lead source, measured to revenue rather than to lead count

Most shops evaluate marketing on lead volume because lead volume is the number the vendor reports. Ask instead for booked revenue and gross profit by lead source, divided by the jobs that source produced. A source generating half the leads of another can easily produce more profit if it brings different work. Getting the source data clean enough to do this is its own project — marketing attribution for home service businesses covers why that column is usually a mess and what to do about it.

Where this goes wrong, and how to catch it

Language models compute arithmetic inconsistently. This is the single most important operational caveat in this entire piece, and it is why the final call stays human.

Before you trust any output, hand-check one row. Pick a job type, pull three actual jobs of that type out of your export, and do the math on a calculator. If the model's number for that category is materially off, the analysis is not usable and you need to either restate the question or move the calculation into a spreadsheet and use AI only to interpret the result. That second approach — spreadsheet does the arithmetic, AI does the reading — is more reliable and is what I would recommend to anyone who does not enjoy verifying.

Four other failure modes worth knowing:

Drive time deserves its own warning. Under the Fair Labor Standards Act, non-exempt field technicians must be paid for travel between job sites during the workday — the Department of Labor's guidance on hours worked is clear on this. That is real labor cost you are already paying. If your analysis only counts on-site hours, every job in your outer service radius is showing a margin you are not actually earning.

What does AI business analytics actually change for a service business?

It changes which questions you can afford to ask. Before, answering "what is my gross profit per truck-hour by job type" meant either building a spreadsheet model yourself or paying someone to. Now it is an afternoon with an export and a careful prompt. The analysis does not decide anything — it narrows what you are deciding about. A shop that learns its highest-revenue category is its fourth-most profitable does not stop doing that work. It reprices it, routes it differently, or stops chasing more of it in its marketing.

The judgment stays yours because the data cannot see what you see. It cannot see that the thin-margin category is what keeps three good techs busy in February. It cannot see that a lead source with mediocre numbers sends the commercial referrals that pay for the year. It cannot see which customer is worth losing money on. Those are the calls only the owner can make, and every analysis in this piece is an input to them rather than a substitute for them. The same goes for your crew: nothing here tells you how to train a tech or whether to promote one. It tells you where to go look.

A realistic first pass, start to finish

  1. Write down two questions you currently answer by feel. Be specific enough that a number would settle them.
  2. Export twelve to eighteen months of completed jobs from your field service software. Twelve months minimum so seasonality does not distort you.
  3. Strip customer names, addresses, and contact information. Keep an anonymous customer ID if you want repeat analysis.
  4. Consolidate your job type column down to a short list that reflects how you actually think about the work, and save the mapping for next time.
  5. Fix or approximate the labor column. Real clocked hours if you have them, one consistent loading factor if you do not. Write down which you used.
  6. Ask one question with full context — what the columns mean, how labor cost was built, what output shape you want, and a note to show median alongside mean.
  7. Hand-verify one row against a calculator before you believe anything.
  8. Make one change based on what you found. One. Repricing a category, rerouting estimates above a threshold, or cutting a lead source.
  9. Re-run the same analysis in ninety days using the identical prompt, so the comparison is apples to apples.

That last step is where the compounding happens and where most people quit. Save the prompt in a note. The value of this practice is not the first answer — it is having a repeatable way to check whether the change you made worked. An owner who runs the same four analyses every quarter for a year knows more about their business than one who buys a dashboard and glances at it.

If this is your first analytics work of any kind, resist the urge to also fix your phones, your reviews, and your scheduling in the same month. The sequencing argument in the 90-day AI implementation roadmap applies directly here, and it is genuinely fine for the answer to be one analysis and nothing else. At Turnkey AI we tend to push owners toward exactly this kind of narrow, verifiable first project rather than a stack of tools nobody has time to learn.

Questions owners actually ask about this

These come up nearly every time.

Frequently asked questions

Do I need to buy an analytics tool to do this?

No. The method in this piece uses a CSV export from the field service software you already run, a spreadsheet to clean it, and a general-purpose AI tool to interpret it. Dedicated analytics products can help once you know which questions matter, but buying one before you know that usually produces a dashboard nobody opens.

How much job history do I need before the numbers mean anything?

Twelve months minimum, so seasonality does not skew the result. Eighteen to twenty-four months is better because it lets you see whether a pattern holds across two cycles. If you have fewer than about 300 completed jobs, treat the output as directional and verify anything surprising by hand.

Is it safe to paste my job data into an AI tool?

Strip customer names, addresses, phone numbers, and email addresses first. None of them are needed to answer margin questions. Keep an anonymous customer ID if you want repeat-purchase analysis, and use tech first names or initials only. Check your AI tool's data retention settings before the first paste.

Why does the AI get different totals when I ask the same question twice?

Language models compute arithmetic inconsistently, especially on large data sets. Always ask the model to state how many rows it processed, and hand-check one category against a calculator. If the numbers do not hold, move the arithmetic into a spreadsheet and use AI only to interpret the result.

What if I do not track labor hours per job?

Use a consistent loaded hourly cost per tech as a stand-in: base wage multiplied by a loading factor covering payroll taxes, insurance, phone, and vehicle. The absolute margin figures will be approximate, but the ranking of job types is usually stable, and the ranking is what drives your decisions.

Should I make several changes at once based on what I find?

Make one. If you reprice a category, reroute estimates, and cut a lead source in the same week, you will not be able to tell which move produced the result ninety days later. One change, one re-run of the identical analysis, then the next change.

About Turnkey AI

Turnkey AI helps service businesses put practical AI tools and automation to work — AI receptionists, automated lead follow-up, scheduling, review requests, and more — so owners reclaim time without adding headcount.