inquiry

Help desk metrics: what is worth watching when one person answers everything

October 7, 2026 ・ Halict Editorial

The standard article on help desk metrics lists seventeen of them. Ticket volume by channel, tickets per agent, agent utilisation, first contact resolution rate, cost per ticket, backlog, customer satisfaction, net promoter score, and a leaderboard. It is a good list for a support organisation of forty people with an analyst maintaining the dashboard. It is close to useless for the two people who answer everything that arrives through a form and also do three other jobs.

The reason is not that small operations do not need numbers. It is that most of that list measures the distribution of work across agents, and where there is nothing to distribute, those measurements produce activity rather than insight. Worse, several of them create incentives that make the service worse: anything measured per agent rewards closing fast, and closing fast is not the same as finishing.

The metrics that assume a team that does not exist

Four groups of metrics can be set aside immediately at small scale, and it saves a great deal of dashboard building to say so plainly.

Anything per agent. Tickets handled per person, replies per person, average handling time per person. With one or two people the figure is the total, and comparing a person against themselves month to month mostly measures how busy the rest of their job was.

Utilisation and occupancy. These exist to schedule shifts in a contact centre. A team that answers the queue between other work has no shift to schedule.

Cost per ticket. The arithmetic requires an allocated salary cost and a stable volume, and at low volume the figure swings so widely that no decision should rest on it. The decision it is meant to inform, whether to hire, is better made from the age of the backlog.

Net promoter score. It measures a relationship with the organisation, not an interaction with the help desk, and at a few dozen responses the confidence interval is wider than any plausible change.

What remains is a short list, and a short list that is actually read beats a dashboard that is not.

Four numbers that work at any size

The age of the oldest item with no reply. One number, in days. It cannot be improved by volume going down, it cannot be averaged away, and it points at a specific request that somebody has to deal with today. Every other metric on a small queue is a refinement of this one.

The share of open items where the team owes the next move. Split the open queue into two piles: waiting on the team, and waiting on somebody else. A queue that is mostly waiting on customers needs better chasing and reminders. A queue that is mostly waiting on the team needs either more time or fewer promises. The same total backlog calls for opposite responses depending on the split, which is why the total on its own is not actionable.

The reopen rate. The proportion of finished items that come back. This is the check on speed, because fast first replies and quick resolutions can both be manufactured by closing things prematurely. When response times improve and reopens climb at the same time, the improvement is not real.

Volume by reason. Not by channel, not by hour. By reason, using a short fixed list of eight to fifteen codes. This is the only metric on the list that reduces future work rather than describing past work, because it names the thing that, once fixed, stops the request arriving at all.

The definitions differ more than the numbers

Two teams both reporting a two-hour average first response may be measuring completely different things, which is why a benchmark from an industry report is worth very little without the definition attached.

First reply time, as Zendesk defines it, is the time between ticket creation and the first public agent comment. An automatic acknowledgement is not a public agent comment, which is the correct treatment: an auto-reply is not an answer. A tool that counts the auto-reply will report a first response time of a few seconds and mean nothing by it.

Requester wait time is a different metric, and the distinction is the useful one. It counts the time a ticket spends in the New, Open and On-hold statuses, which means it excludes time spent waiting for the customer. Two numbers, one including customer delay and one excluding it, answer two different questions: how long the customer experienced, and how long the team took.

One-touch resolution has the most surprising definition. Zendesk counts a ticket as one-touch when it is solved or closed with a single agent reply or no replies at all. Agent replies only count after the conversation has transferred to the agent, a customer replying to say thank you does not break the count, and a customer reopening with another question that gets another answer does.

Business hours versus calendar hours changes every duration metric, and the two are not interchangeable across metrics either. Status-based metrics such as requester wait time capture the schedule in force at each status change, while single-event metrics such as first reply time apply one schedule to the whole ticket. That is usually the answer when two reports disagree by exactly one working day.

Averages lie at low volume

Thirty items a week is not enough data for an average to be stable. One request that sat over a holiday weekend moves the monthly mean by hours, and the natural response to a mean that jumps around is to stop believing it.

Instead of Use Because
Average first response time Ninetieth percentile, plus the single worst case The tail is where complaints come from
Average resolution time Count of items older than a threshold A count is a to-do list, an average is a statistic
Backlog total Backlog split by who owes the next move The split determines what to do
Satisfaction percentage The text of every comment received Six responses cannot support a percentage
Tickets per person Total volume by reason Reasons can be reduced, people cannot be cloned

The right-hand column has a property the left-hand column lacks: each one names an action. That is the test for a metric on a small team. If a number changes and nobody knows what to do differently, it is decoration.

Percentiles are worth the small effort they take. Sorting ninety days of response times and reading the value nine tenths of the way down the list is a spreadsheet operation, and it produces a number that describes the experience of the unlucky customer rather than the typical one.

Satisfaction scores need care below a few hundred responses

Satisfaction surveys are the metric most often adopted too early, because the mechanics are easy and the output looks authoritative.

Start with what the score is made of. Zendesk maps survey responses to good and bad, and the mapping depends on the scale: on a one to five scale, one to three count as bad and four and five count as good. A steady stream of fours therefore reads as unqualified success. On a one to three scale, one and two are bad. Knowing the mapping is necessary before quoting a percentage to anybody. Viewing the results is also a paid-tier feature, requiring Growth or above on Zendesk Suite, or Professional or above on Zendesk Support.

Then consider the sample. Thirty finished items a week with a twenty per cent response rate produces six responses. One unhappy customer moves that from 100 per cent to 83 per cent, and nothing whatever has changed about the service. Trending that percentage month over month produces a chart of noise.

The useful practice at this size is to collect the survey and read the comments rather than the score. Six sentences written by real customers contain more actionable information than the percentage they roll up into, and the sentences point at a reason code, which is the metric that reduces future work. Once volume reaches a few hundred responses a quarter, the percentage starts to mean something and can be trended.

Getting the numbers without a dashboard

Small operations rarely have analytics, and they do not need it. Four fields per record are enough.

The time it arrived. The time the first outbound reply was sent. The owner. The current stage, and the stage history if the tool records it. A reason code is the fifth and most valuable, and it is the one that has to be filled in by a person.

With those, a single export into a spreadsheet answers every question on the short list. Received time and first reply time give the percentile. Stage gives the split between waiting on the team and waiting on others. Stage history gives the reopen rate. The reason code gives the pivot table that decides what to fix.

Two practical cautions. The export has to include the timestamps, not just the current state, or response time cannot be calculated after the fact, so this is worth checking before the numbers are needed rather than after. And the reason code has to be an internal field with a fixed list of options, not free text, because free text cannot be counted. Tools that keep internal fields alongside the response and push each new response into a spreadsheet row make this a five-minute job each week rather than a project. The use cases where this pays off fastest are the repetitive ones, where the same three reasons account for most of the volume.

The review that makes the numbers matter

Fifteen minutes, on a named day, by one person. Three questions, in order.

What is the oldest item with no reply, and what happens to it today? This is the only question with a mandatory action attached, and answering it is most of the value of the whole exercise.

Which of the open items are waiting on the team rather than on somebody else, and is that number growing? A number that grows for three weeks running is the signal to change the promise or the staffing, and it arrives well before the complaints do.

What were the top three reasons this week, and is any of them fixable at the source? One fix a month to a form, a help page, a confirmation email or a price list compounds. This is the only line of work that makes the queue smaller rather than faster.

Anything that cannot be answered in fifteen minutes from one export does not belong in the review yet.

What to change first

Stop reporting the average, and put one number where the team can see it: the age of the oldest item with no reply. Add a reason code with a fixed list of options to the intake form this week, since the reason column is what turns a queue into a list of things to fix, and check that whatever holds the queue can export both the timestamps and that code, as the Halict sample does.

Q1. What is the single most useful help desk metric for a small team?

The age of the oldest item that has had no reply. It is one number, it cannot be flattered by volume changes, and it always names a specific request that somebody has to act on today. Averages describe the past; this one produces a decision.

Q2. Does an automatic acknowledgement count as a first response?

It should not. Zendesk measures first reply time to the first public agent comment, which excludes automated replies, and that is the right treatment because an acknowledgement is not an answer. If a tool counts the auto-reply, the reported response time will look excellent and mean nothing.

Q3. Is it worth running a satisfaction survey on a small queue?

Yes, but for the comments rather than the score. At a few responses a week the percentage moves wildly on a single rating, while the written comments point directly at a cause worth fixing. Trend the percentage only once there are a few hundred responses to average.

Q4. How can response times be measured without reporting tools?

Export the queue with the arrival time and the time of the first outbound reply, then sort and read the value ninety per cent of the way down the list. That percentile, plus the single worst case, describes the service better than a mean. The one requirement is that the export contains timestamps rather than only the current state.

Q5. Why measure volume by reason rather than by channel?

Because a reason can be eliminated and a channel cannot. Knowing that 40 per cent of requests arrive by email is background information. Knowing that 40 per cent ask where an order is points at a tracking email that is missing or unclear, and fixing that removes the work instead of speeding it up.

All guides

Help desk metrics: what is worth watching when one person answers everything | Halict