A 90-minute job booked into a 60-minute slot does not fail at minute 61. It fails at 2:40 in the afternoon, when your third customer of the day is at the front window wondering where the truck is. AI for job duration estimates scheduling tools exist to fix that problem. They replace the slot lengths someone typed into your system years ago with how long your crews actually take, measured from your own completed jobs.
A service business in the $1M-$5M range has two realistic ways to do this. The first is a percentile lookup table: you build it from exported job history and refresh it on a schedule. The second is a predictive duration model that estimates each job's length from its details. It can live inside your field service software or sit on top of it. I have seen owners pick the wrong one in both directions. This piece compares them on what matters: the data each needs, how long before you can trust it, and how each one fails.
Why are job duration estimates wrong in the first place?
Job duration estimates run short for a predictable reason: people underestimate how long their own tasks take, even tasks they have done many times. Psychologists call this the planning fallacy. In the original study by Buehler, Griffin, and Ross, published in the Journal of Personality and Social Psychology, students predicted their honors theses would take 33.9 days on average. They actually took 55.5 days. Your 60-minute tune-up slot was probably set the same way, by someone who remembered the clean jobs and forgot the seized disconnect, the clogged condensate line, and the customer who wanted to talk about their thermostat for twenty minutes.
Three things make it worse in the trades:
- The slot was set once. Most shops enter default durations in Jobber, Housecall Pro, or ServiceTitan during setup and never revisit them, even after the equipment mix, the crew, and the code requirements change.
- Overruns compound and underruns don't. A job that finishes 20 minutes early rarely pulls the next customer forward, because that customer isn't home yet. A job that runs 30 minutes long pushes back every stop after it. An estimate that is right on average still breaks the afternoon.
- Regulation adds minutes that old history doesn't show. In pre-1978 homes, the EPA Renovation, Repair and Painting rule requires lead-safe practices once a job disturbs more than 6 square feet of painted surface per interior room or 20 square feet outside. Cutting a wall for a line set or a new circuit can cross that line. Containment and cleaning verification take time, and a shop-wide average for that job type hides it.
Wondering which AI tool pays for itself first? Request a free AI audit and find out.
What job data do you need before any duration estimate is trustworthy?
A duration estimate can only be as good as the timestamps behind it. Before you compare tools, check whether your job history includes these fields:
- A consistent job type. "Tune-up," "AC maint," and "spring check" have to collapse into one code. Free-text job titles are the most common reason duration data is useless. This cleanup is also one of the few places where AI earns its keep right away: a chat model can sort a few thousand free-text descriptions into fifteen job types in minutes, and you check the result.
- Real start and finish taps. Housecall Pro records On My Way, Start, and Finish. ServiceTitan records Dispatched, Working, and Done, plus Hold for parts runs. Jobber tracks visits and timesheets. The time that matters is from Start to Finish, which I'll call wrench time. Drive time belongs to routing, not to the job.
- The tech or crew. A second-year tech and a 20-year tech will not run the same job in the same time, and you shouldn't expect them to.
- The attributes that change difficulty. Equipment location (attic, closet, garage, crawlspace), home age, number of stories, square footage or lot size, first visit or repeat customer. Pick the three that cause the most variance in your trade. Don't try to capture twenty.
- Outcome flags. Did the job need a parts run, a return trip, or an add-on sold on site? If a tech sold and installed a capacitor during a maintenance visit, the job got longer because it got bigger. That is an add-on, not an overrun. Without this flag, your best salespeople look like your slowest techs.
Then filter the obvious junk. Every shop has an end-of-day tapper who marks all four jobs Finished from the shop parking lot at 5:45. Drop durations under about a quarter of the median and over about four times the median, or flag them for review. Don't let them quietly feed the average. If more than about one job in five has unusable timestamps, fix the tapping habit before you buy anything. Some shops require a Finish tap before an invoice can be sent, which solves most of the problem in a week.
Option one: the percentile lookup table
The lookup table is the less impressive approach and usually the right first move. Export six to twelve months of completed jobs, filter out the bad records, and calculate three numbers for each job type: the job count, the median wrench time, and the 80th percentile wrench time (the time that 80 percent of jobs finish within). A spreadsheet's PERCENTILE function handles this. A chat assistant can write the formulas and summarize the results if you remove customer names and addresses from the export first.
Then make one booking rule: book early-day jobs at the 75th to 80th percentile, and the last job of the day at the median. Morning slots protect the afternoon, so they get the safer number. The last job has nothing behind it, so a tighter estimate costs you little if it runs long. That one rule fixes most cascade problems without any software change. You just update the default durations in your field service platform.
Where it is strong: you can see every number and argue with it. A dispatcher understands it in five minutes. It works with a few hundred jobs. Where it is weak: it treats every job inside a type as the same. An attic furnace and a garage furnace get the same slot unless you split them into separate job types. You also have to rebuild the table on a schedule, or it goes stale the same way the original defaults did.
Option two: a predictive duration model
A predictive model estimates each job individually. It considers the job type, the attributes (attic, 1960s house, two stories), the assigned tech, and the season, and predicts a duration for that specific combination. This is what most vendors mean by learned or AI-predicted job durations. Some dispatch platforms build it in. Others offer it through an analytics layer that reads your job history through an integration. If you go the second route, check which AI tools actually connect to Jobber, Housecall Pro, and ServiceTitan before you assume the data will flow.
A model can capture combinations a table misses. For example, a particular tech might be fast on rooftop units but slow on crawlspace work. The cost is that it needs much more data, and it fails less visibly. When a table is wrong, you can see the wrong number. When a model is wrong, you see a confident suggestion.
Before you turn one on, get answers to these questions from the vendor:
- Are durations learned from my completed jobs, or from industry defaults mixed with mine?
- How many completed jobs of a type does the model need before it overrides my default?
- Can a dispatcher see why it predicted 2 hours 10 minutes, meaning which attributes moved the number?
- Does it separate drive time from wrench time?
- When a dispatcher overrides a prediction, is the override saved so I can audit it later?
How do the two job duration approaches compare side by side?
| Dimension | Percentile lookup table | Predictive duration model |
|---|---|---|
| Data needed to start | Job type plus start and finish times | Job type, times, tech, and 3-5 consistent attributes |
| Time until reliable | Weeks, if you have ~30 clean jobs per type | Usually a full year, to capture seasonality |
| Handles attic vs. garage, old vs. new home | Only if you split them into separate job types | Yes, if those fields are filled in consistently |
| Handles differences between techs | With a manual per-tech multiplier | Built in |
| Explainability | Total: you can see every number | Varies by vendor; ask to see the drivers |
| Upkeep | Monthly or quarterly rebuild | Retrains on its own; needs periodic audits |
| How it fails | Goes stale quietly if nobody refreshes it | Confidently wrong when input data is messy |
| Best fit | 2-10 trucks, consistent job menu | Higher volume, varied jobs, clean history |
How long before AI duration estimates are reliable?
These thresholds are rules of thumb, not guarantees. For a lookup table, about 30 clean completed jobs per type gives you a median that stops moving much from month to month. The 80th percentile is noisier and needs closer to 50. Job types you run four times a year will never have enough data. Keep a manual default for those and move on.
A predictive model needs several hundred jobs per job family. More importantly, it needs a full cycle of seasons. Summer attic changeouts, first spring mows on overgrown lots, and January no-heat calls in old houses all behave differently. A model trained on six months has never seen half the year. Some trades have a built-in split: a pest control initial service runs longer than the recurring quarterly visit at the same property, and a first deep clean runs longer than the recurring clean. Make those separate job types in either approach, or the average will be wrong for both.
Also restart the clock after any real change in the work. The refrigerant transition is the clearest example. The EPA's Section 608 program already requires refrigerant recovery on service and disposal, and recovery time grows with system size and outdoor heat. New residential systems now use A2L refrigerants such as R-454B and R-32, with their own handling procedures. Install durations recorded on R-410A equipment before 2025 don't describe current installs. Treat them as a starting guess, not as training data.
Whichever approach you choose, check it the same way. Take last month's completed jobs, compare the predicted durations with the actual ones, and calculate your slot hit rate: the share of jobs that finished inside their booked slot. Track it weekly. If the 80th-percentile rule is working, the hit rate should land near 80 percent. If it is far below that, your data has a problem the tool can't fix.
Worked example: the water heater swap that was never a two-hour job
The numbers here are illustrative, but the pattern is very common in plumbing shops. A shop books every 50-gallon gas water heater replacement as a two-hour job. The installers complain, the afternoon calls run late, and the owner assumes the crew is slow.
The export shows 40 swaps over the year, with a median of 2 hours 35 minutes and an 80th percentile of 3 hours 20 minutes. The two-hour slot was short on almost every job. The bigger finding comes from splitting by location. Garage swaps have a median of about 2 hours 15 minutes. Attic and closet swaps have a median near 3 hours 30 minutes, because of the drain pan, the condensate or relief line routing, the tight access, and the trip up and down the ladder with a 120-pound tank.
Code makes the gap bigger. Installing a required thermal expansion tank on a closed system (IPC 607.3 where the IPC is adopted) adds time. In California, Health and Safety Code 19211 requires seismic strapping. Neither shows up in history recorded before the shop started doing them consistently.
The fix didn't require a model. The shop created two job types, garage swap and attic/closet swap. The CSR started asking one question on intake: "Where is the water heater?" That question is also how you stop losing information before a job reaches the calendar. Garage swaps now get a 2.5-hour slot. Attic swaps get a half-day and a second person. The installers aren't any faster. The schedule just finally matches the work.
Which job duration approach should you pick?
Pick the percentile lookup table when: you run roughly 2-10 trucks, your job menu is fairly stable, your timestamps are only partly clean, and you want a fix running this month. That describes most $1M-$5M shops.
Pick a predictive model when: you have high job volume, the same job type varies a lot because of property and equipment details, you have at least a year of clean Start/Finish history, and someone on your team will audit its predictions every month. It works best if your field service platform already offers it trained on your own jobs.
For most owners the order is the table first, then possibly a model. Building the table forces the cleanup any model needs anyway: consistent job types, honest timestamps, add-ons flagged. A year of that discipline gives you good model training data, and you may find the table is good enough. If you want to see how duration fits into the rest of the board, this run-today audit for scheduling and dispatch covers the other steps, and our AI scheduling page explains how we approach it with shops.
What should stay with the dispatcher, not the AI scheduling tool?
An AI scheduling tool predicts durations from history. It can't know what your dispatcher heard on the phone. When the customer says the unit "has been making that noise for three weeks," an experienced dispatcher adds time. A new tech's first solo install gets a buffer. The property with the locked side gate and the dog everyone remembers gets extra minutes. These are judgment calls. The tool should suggest a number, the dispatcher should make the decision, and every override should be saved so you can see whether the human or the model was right more often.
Duration data also has value beyond scheduling. If a flat-rate job keeps running 50 percent over its booked time, your scheduling problem is also a pricing problem. Duration is one of the clearest margin leaks you can find with AI, and it is worth reviewing with whoever builds your flat-rate book. Faster, more accurate quotes depend on the same history, which is why AI quoting for contractors starts from the same job records.
What do owners ask about AI job duration estimates?
Should I just book every job at its average time?
No. An average-length slot means roughly half your jobs run over, and overruns stack through the day. Book morning and midday jobs at the 75th to 80th percentile and the last job of the day near the median.
My techs forget to tap Start and Finish. Is my data useless?
It isn't useless, but it needs filtering. Drop durations that are obviously impossible. If fleet GPS is available, use arrival data to fill gaps. Then make the Finish tap required before an invoice can go out. Within a month you'll have usable history going forward.
Do I need separate durations for each technician?
Not at first. Get the job-type table stable before anything else. Once it is, a simple per-tech multiplier (one tech runs about 1.2 times the shop median on installs, for example) covers most of the difference. Only a predictive model handles tech differences automatically.
How often should job duration estimates be updated?
Rebuild a lookup table monthly during your busy season and quarterly otherwise. Also rebuild right away after any change to equipment, code requirements, refrigerant, or crew makeup.
Can ChatGPT build the duration table for me?
Yes. It can write the spreadsheet formulas, sort free-text job titles into consistent types, and flag outliers. Remove customer names, phone numbers, and addresses from the export first, and spot-check its percentile results against a few jobs you remember.
Frequently asked questions
Should I just book every job at its average time?
No. An average-length slot means roughly half your jobs run over, and overruns stack through the day. Book morning and midday jobs at the 75th to 80th percentile and the last job of the day near the median.
My techs forget to tap Start and Finish. Is my data useless?
It isn't useless, but it needs filtering. Drop impossible durations, use fleet GPS arrival data to fill gaps if you have it, and require a Finish tap before an invoice can go out. Within a month you'll have usable history going forward.
Do I need separate durations for each technician?
Not at first. Get the job-type table stable first, then add a simple per-tech multiplier. Only a predictive model handles tech differences automatically.
How often should job duration estimates be updated?
Rebuild a lookup table monthly during your busy season and quarterly otherwise. Also rebuild right away after any change to equipment, code requirements, refrigerant, or crew makeup.
Can ChatGPT build the duration table for me?
Yes. It can write the formulas, sort free-text job titles into consistent types, and flag outliers. Remove customer names, phone numbers, and addresses from the export first, and spot-check its results against jobs you remember.