First response time is the gap between something arriving and a person answering it. It is the most quoted measure in any inbox, queue or intake process, and it is also the one most often reported without anybody agreeing what it counts. Two people can look at the same month of enquiries and produce first response figures that differ by a factor of five, both correct, because they started the clock in different places.
That matters more than a reporting quibble. A first response figure is usually the number a target gets attached to, and a target attached to a badly defined measure produces behaviour that improves the measure and not the service. Sending a holding message to everything within minutes will do wonders for the number. This is what the measure actually contains, how to read it honestly, and which changes move it.
What the number is measuring
First response time starts when a request arrives and ends when a person sends the first substantive reply. Everything difficult is in the word substantive, and in the word arrives.
Distinguish it from two neighbours. Resolution time measures arrival to the request being finished, which is a different question and usually driven by factors outside the responder's control. Time to acknowledge measures arrival to any confirmation at all, including an automated one. That third measure is worth tracking separately, and it is worth keeping out of the first response figure, because merging them makes an automated receipt look like an answer.
The reason to care about the first number specifically is that it is the point at which the requester stops guessing. Someone who has submitted an enquiry and heard nothing has no way to tell whether it arrived, whether anyone owns it, or whether following up would help. Most duplicate submissions and most chasing messages come from that uncertainty rather than from impatience about the answer itself. The first substantive reply ends the uncertainty, which is why a slow first response generates extra work rather than simply postponing it.
There is a second reason. First response time is the earliest measure available. A resolution figure for this month is not knowable until the month's work is finished, while a first response figure is knowable within a day. As an operational signal it arrives in time to act on, and that is most of its value.
The four decisions inside the number
Before comparing any two first response figures, settle these four. Each one changes the result substantially.
Where the clock starts. The moment the submission was made, or the moment it appeared in the place the team works from? Those differ when a form writes to a spreadsheet that somebody checks twice a day, or when an email passes through a rule that files it. Starting the clock at arrival in the team's view flatters the number by hiding the queue before it.
Whether the clock runs overnight. Calendar hours count the whole elapsed time. Business hours count only the hours the team was open. A request that arrives at 5:30pm on Friday and gets answered at 9:15am on Monday is either 63 hours or 45 minutes, depending only on the setting. Both are defensible. Neither is comparable to the other, so the choice has to be stated wherever the figure is.
What counts as the response. An automated receipt, a holding message written by a person, or an answer that engages with the content. A definition that admits holding messages makes the measure trivially improvable and worthless. The workable rule is that a reply counts when it either answers the request, asks a question needed to answer it, or states a specific commitment such as a named owner and a date.
Which requests are in scope. Spam, duplicates, internal test submissions and messages that turn out to be replies to an existing thread all distort the figure if included. Decide the exclusions in advance and write them down, because deciding them while looking at a bad month is how reporting loses credibility.
Why the average is the wrong summary
The average first response time is the figure most systems display and the least useful one to act on. The distribution of response times is heavily skewed, and averages assume it is not.
Consider ten requests. Nine are answered within an hour. One sits over a long weekend and is answered after 40 hours. The average is 4.9 hours, which describes none of the eleven. Reported to a manager, that average looks like a mediocre month. Reported as a median, the same month reads as one hour, which describes nine of the ten and hides the one that actually failed somebody.
| Summary | What it tells you | Where it misleads |
|---|---|---|
| Average | The total load divided by the count | One long case drags it; no request may resemble it |
| Median | The typical experience | Says nothing about how bad the tail gets |
| 90th percentile | How bad it gets for the unlucky tenth | Needs enough volume to be stable |
| Slowest case | The single worst experience given | Can be a one off, so read it as a case, not a trend |
| Count over threshold | How many exceeded the target | Requires a target to already be set |
The pair worth reporting together is the median and the count over threshold. The median describes the normal day. The count over threshold names the failures as cases rather than as a shift in a decimal. Improvement work then has an actual queue to work through, since each item in that count is a specific request with a specific reason it sat.
For small volumes, the honest answer is to stop summarising. Under about thirty requests in a period, list them. The reasons the slow ones were slow will be visible in the list and invisible in any statistic drawn from it.
Why the figure slips without anyone working slower
First response time usually degrades for structural reasons rather than because effort dropped. Four causes account for most of it.
Arrivals are not evenly spread. Requests cluster, around deadlines, around announcements, around Monday mornings. A team sized for the average day has a backlog on the clustered day, and the backlog pushes first response times for everything behind it.
Nothing is owned until somebody claims it. In a shared inbox, a new arrival is everybody's and therefore nobody's. The most reliable predictor of a slow first response is a request that nobody has been made responsible for. Time spent unowned is invisible in most reporting, because the clock is running while no work is happening and no one is aware of it.
The queue is not visible as a queue. A list sorted by arrival date shows what came in. It does not show what has been waiting longest without a reply, which is a different sort and the only one that helps. Without that view, the oldest untouched item is the one least likely to be seen.
Some requests cannot be answered by whoever picks them up. These need routing, and routing is where hours disappear. A request that passes through three people before reaching the one who can answer it accumulates its wait in the handovers, not in the answering.
None of these are solved by asking people to reply faster. Three of the four are answered by making ownership and waiting time visible, and the fourth by deciding routing rules in advance rather than case by case.
Setting a target that is not a guess
A target picked because it sounds impressive gets ignored by the second week. Build it from three inputs instead.
Start with what the current distribution says. Look at the median and the 90th percentile for the last few months, in business hours, with the exclusions applied. A target set slightly inside the current 90th percentile is demanding and reachable. A target set at the current median means half the work already fails it on arrival.
Then check it against what the requester needs. Different intake types have genuinely different clocks. An application with a closing date needs a response before the date, and a response in four hours has no more value than one in two days. A fault report from someone who cannot work has a clock measured in hours. A general enquiry sits between them. One target across all of them either over serves the patient cases or fails the urgent ones.
Then check it against capacity on the worst day rather than the average one. A target that holds on Tuesday and breaks every Monday is a target the team learns to treat as advisory.
State the target with its definition attached, every time. "Median first response under four business hours, excluding spam and duplicates, clock starting at submission" is a target that can be audited. "Four hour response time" is a slogan, and six months later nobody will remember which of the four decisions it assumed.
What actually moves it
In rough order of effect, and none of these are about typing faster.
Assign an owner at arrival. Whether by rotation, by category or by whoever is on duty, the aim is that no request exists for more than a few minutes without one person responsible for it. This single change moves the tail more than anything else, because the tail is made of unowned items.
Make waiting time the default sort. The working view should show the longest wait without a reply at the top, not the newest arrival. The change is small and the effect is immediate, because the items that damage the figure become the items that are visible.
Separate the acknowledgement from the response. An automatic receipt that states what was received and when a human reply can be expected removes the requester's uncertainty and with it most of the chasing. It does not count as the first response and should not be allowed to, but it buys real time and reduces duplicate submissions.
Ask for what is needed on the form. A large share of slow first responses are slow because the first reply had to ask for information the request should have carried. Fixing the intake fields converts a two message exchange into one, which improves both the first response figure and the resolution figure at the same time.
Prepare the replies that recur. Most intake queues have a handful of situations that make up the majority of arrivals. Saved wording for those, edited rather than composed, cuts minutes off each one. The gain is in the fifty ordinary cases, not the unusual one.
These depend on the request carrying an owner, a status and its own history rather than existing as a row in a sheet or a message in a shared mailbox. The shapes that need this are set out in use cases, and how owners, statuses and history behave per response is described in features.
Measuring it when everything arrives in different places
A single first response figure requires a single place where the clock is readable. Most teams do not have that, because forms land in one place, email in another, and chat in a third.
Two workable answers exist. Bring the intake together, so that submissions from every route become records of the same kind with an arrival timestamp and a first reply timestamp. Or measure the routes separately and stop reporting a combined number, since a combined figure assembled by hand from three sources will be wrong in a different way each month.
What does not work is estimating. A first response figure reconstructed from memory or from a spreadsheet column that people fill in when they remember is not a measurement, and it will always be more flattering than reality. Either the timestamps are recorded automatically or the number should not be reported.
What to change first
Write down the four definitions, then switch the working view to sort by longest wait without a reply and give every arrival an owner within minutes. Those two changes attack the tail, which is where the damage is, and the median will follow. If arrivals currently land in a sheet or a shared mailbox where neither the owner nor the waiting time is visible, that is the thing to replace, and the Halict demo shows what it looks like when both sit on the response itself.
Q1. How is first response time calculated?
It is the elapsed time from a request arriving to the first substantive reply from a person. Four settings change the result: whether the clock starts at submission or at appearance in the team's view, whether it runs outside business hours, whether automated receipts and holding messages count, and which requests are excluded. State all four alongside any figure, or it cannot be compared to anything.
Q2. Should automated replies count as the first response?
No. Counting them makes the measure trivially improvable and hides the real wait. Track acknowledgement separately, and count a reply as the first response only when it answers the request, asks a question needed to answer it, or commits to a named owner and a date.
Q3. What is a good first response time?
There is no single figure, because the requester's clock differs by intake type. An application with a closing date, a fault report from someone unable to work, and a general enquiry have genuinely different needs. Build the target from the current 90th percentile in business hours, check it against what each intake type requires, and test it against the busiest day rather than the average one.
Q4. Why is average first response time misleading?
Response times are heavily skewed, so one long case moves the average away from every actual case. Nine replies within an hour plus one after 40 hours average 4.9 hours, which describes none of them. Report the median together with the count of requests that exceeded the target, so the failures appear as specific cases rather than as a decimal.
Q5. Should first response time be measured in business hours or calendar hours?
Either, as long as it is stated. A request arriving at 5:30pm Friday and answered at 9:15am Monday is 63 calendar hours or 45 business minutes. Business hours describe team performance, calendar hours describe the requester's experience, and mixing the two between reports is what makes trends meaningless.
Q6. What is the fastest way to improve first response time?
Give every arrival an owner within minutes and sort the working view by longest wait without a reply. Unowned time is where the slow cases come from, and sorting by arrival date hides exactly the items that need attention. Both changes need the owner and the waiting time to live on the individual request rather than on a file.
