What to Measure in the Ninety Days After Go-Live
· 11 min read · Faceela
Six weeks after go-live, a steering committee asks the only question it has ever really wanted answered: is it working?
What comes back is a slide with three figures on it. Ninety-four per cent of users have logged in. Training completion is at ninety-nine per cent. System availability is 99.8 per cent. Everyone nods, the meeting moves on, and four months later the finance director quietly mentions that the close is still taking eleven days and the sales team has gone back to quoting in a spreadsheet.
Every number on that slide was true. Not one of them was about the business.
This is the most common failure of the post-go-live period, and it is not a measurement problem so much as a courage problem. The metrics that get reported are the ones that are easy to produce and safe to present. The ones that would tell you something are harder to produce and, in the first two months, frequently embarrassing.
The rule: measure the work, not the software
A useful post-go-live measure has three properties. It describes something the business did rather than something the system did. It existed as a concept before the project started, so it can be compared to something. And it is owned by a person whose job it is, not by the project.
Apply that test to the usual reporting and most of it disappears.
Logins and active users measure whether people can get in, which stopped being interesting on day three. A user can log in every morning and still be running the business from a spreadsheet.
Records created measures typing. A system full of half-completed records that nobody trusts scores brilliantly on this.
Training completion measures attendance. The relationship between attending a session before go-live and being able to do the job after it is weak enough to be nearly random, which is why the training that changes behaviour happens four to six weeks after go-live rather than before.
Uptime matters, and it belongs in an IT report, not a business one. Nobody has ever abandoned a system because it was down for forty minutes. They abandon it because raising a delivery note takes four screens.
Tickets closed is the worst of the group, because it can be improved by closing tickets. Ticket volume and ticket ageing carry information. Closure rate carries a target.
"Adoption" as a single percentage is the most dangerous number in this period, because it is usually a composite of the above, and because it invites the conclusion that the project succeeded when the underlying processes have quietly degraded. It is the same class of error as a dashboard that is technically accurate and still lying to you.
The small set that matters
Seven measures. Not a scorecard of thirty. Seven, because a steering committee will genuinely look at seven and will genuinely not look at thirty, and because each of these has a named owner outside the project.
| Measure | What it actually tells you | Where it comes from | Expected shape in ninety days |
|---|---|---|---|
| Days to close the month | Whether finance can operate the system rather than fight it | Date of the final approved trial balance minus period end | Worse for one to two cycles, then better than before by month three |
| Order to invoice cycle time | Whether the operational chain from order to cash actually flows | Median elapsed days, order confirmation to invoice issue | Worse in weeks one to three, recovering by week six |
| Stock accuracy | Whether the system's picture of the warehouse is believed | Cycle count variance, by value and by line, weekly | Should hold or improve immediately; if it degrades, transactions are not being posted |
| Quotation turnaround | Whether the front end got slower, which is what customers feel | Median hours, enquiry received to quotation issued | Worse in week one, back to baseline by week four |
| Open exceptions ageing | The leading indicator for everything else | Count and age of unmatched receipts, unposted entries, blocked documents, stale drafts | Rises, peaks around week three, then falls steadily |
| Support tickets by process | Where the design or the training is wrong, specifically | Ticket taxonomy by business process, not by module | Halves by week six; anything that does not halve is a design problem |
| Manual workarounds in use | Whether the system is being used or worked around | A weekly count from floor observation, not a survey | Should fall to near zero; each survivor needs a decision |
The last one is the one nobody measures and the one that predicts the following year. A spreadsheet that survives ninety days becomes permanent, and permanent workarounds are how a company ends up with an expensive system and a shadow system, both partly right.
A few notes on the ones that are commonly done badly.
Days to close must be measured to the same definition it had before. If the old close was declared finished when management accounts were circulated, measure to that, not to the point where the trial balance stopped moving. Half of all reported closing improvements are definition changes.
Cycle times should be medians, not averages. One order that sat for forty days because of a customer dispute will drag an average and tell you nothing. The median tells you what a normal transaction experiences, which is what you are trying to fix.
Stock accuracy should be measured by line as well as by value. Value accuracy is dominated by a few expensive items and can look excellent while the warehouse cannot find anything. Line accuracy is what the picker experiences.
Ticket taxonomy by process is the whole point. "Forty tickets in Inventory" tells you nothing. "Fourteen tickets on goods receipt against partially delivered purchase orders" tells you exactly where to spend Tuesday. Categorise by what the person was trying to do, and by cause: defect, design gap, training gap, or data. That four-way split, published weekly, is more useful than any dashboard, because each cause has a different owner and a different remedy. Defects go to the partner. Design gaps go to the process owner. Training gaps go to a session. Data goes to whoever owns the master record — and if that last one is blank, you have found your real problem.
The baseline you should have captured and probably did not
Every one of those measures is a comparison, and comparisons need a before. The before is almost never captured, because in the three months prior to go-live nobody has spare capacity and measuring the old system feels like effort spent on something you are about to throw away.
So you arrive in week six with numbers and no context, and the conversation becomes an argument about whether eight days to close is good. It cannot be answered without knowing what it was.
If you are reading this before go-live, capture six things and it will take a person about two days: days to close for the last three periods, median order-to-invoice days from a sample of a hundred orders, the last stock count variance by value and by line, median quotation turnaround from a sample of fifty, the count of open exceptions of each kind on a stated date, and a written list of the spreadsheets currently in use with what each one does. Store it somewhere that is not the project drive.
If you are reading this after go-live, reconstruct it, and be honest about the reconstruction. It is entirely possible to recover a usable baseline from evidence that already exists.
Closing dates are in email: the message that circulated management accounts has a timestamp, and three of those give you a defensible figure. Order-to-invoice can be reconstructed from the old system's document dates for a sample, even where the transactional detail did not migrate. Quotation turnaround is usually recoverable from the sent folder of whoever issued them. Stock accuracy is in the last count file. Exception counts are the hardest and are often recoverable from the migration cutover pack, because somebody counted open documents in order to move them.
Write the reconstructed baseline as a range rather than a figure, state the method, and get it agreed before you report against it. A range agreed in advance is worth far more than a precise number produced afterwards, because the precise number will be disputed by whoever does not like the comparison, and a dispute about methodology in week eight consumes the credibility the project needs.
There is one thing worth measuring that has no baseline and does not need one: the cost of the exceptions themselves. If you have never quantified what the manual chasing, re-keying and reconciling was costing before, the cost of chaos calculator produces a defensible order of magnitude, and it is more useful as a before-and-after frame than as a business case.
How to read a number that gets worse
Almost everything gets worse first. This is not a comforting platitude; it is a mechanical consequence of replacing practised behaviour with unpractised behaviour, and of a system that now records exceptions that the old one simply absorbed.
The skill is telling the three shapes apart, and the three shapes are distinguishable within about six weeks if you are measuring weekly rather than monthly.
Worse, then better. The number degrades for two to four weeks, flattens, and improves. This is learning, and it needs nothing but patience and the second round of training. Order-to-invoice cycle time nearly always does this. So does quotation turnaround.
Worse, then flat. The number degrades and stays there. This is not learning; it is design. Somewhere in the process there is a step that is genuinely slower than the old way, and no amount of familiarity removes it. Four approval levels where there were two. A mandatory field nobody can populate at the point it is asked for. A document that now requires information from a person who is not in the building. This shape is the most valuable finding of the whole ninety days, and it is the one most often misread as "people need more time".
Better immediately. Treat this with suspicion. A close that got faster in the first month usually got faster because something is not being reconciled. Stock accuracy that improves the week after go-live sometimes means adjustments are being posted to make the count agree rather than the count being investigated. Ask what stopped happening.
The diagnostic that separates the first two shapes reliably is not the metric itself, it is the ticket taxonomy sitting beside it. If cycle time is flat and the tickets on that process are falling, the problem is design. If cycle time is flat and the tickets are also flat, the problem is training. If both are falling and the metric is still bad, you are looking at a data problem — usually a master record that forces a workaround on every transaction.
Exceptions ageing is the number to watch weekly
If a steering committee will only look at one thing, make it this.
Open exceptions — unmatched goods receipts, unposted entries, documents blocked at approval, drafts older than a stated number of days, failed integration messages — are the leading indicator for every other measure in the list. They rise in the first weeks because that is what a new system does, and the shape of the fall is the truest statement about whether the organisation has absorbed the change.
A pile that peaks in week three and declines steadily is a project working normally. A pile that plateaus is a process with no owner. A pile that grows through week six is an emergency wearing the costume of a backlog, and it will surface as a wrong month-end, a delayed customer invoice, or a stock figure nobody trusts.
Report it as a count and an age together. A hundred exceptions all raised this week is a busy week. Thirty exceptions averaging forty days old is an unowned process, and the thirty is the more dangerous number.
The rhythm, and who owns it
The measurement itself has to belong to somebody outside the project team, because a project team measuring its own outcome will, without any dishonesty, choose the definitions that flatter it.
The pattern that works in a mid-market company is unremarkable. Weekly for the first six weeks: exceptions ageing, tickets by process and cause, and the workaround count, reviewed in a thirty-minute standing meeting with the process owners rather than the project. Monthly at thirty, sixty and ninety days: the full set of seven against the baseline, with each measure presented by its business owner rather than by the project manager. The finance director presents days to close. The operations manager presents stock accuracy. The sales manager presents quotation turnaround.
That single change — the owner presents their own number — does more for the honesty of the reporting than any amount of methodology. A project manager presenting somebody else's metric will explain it. An owner presenting their own will fix it.
At ninety days, three decisions are due and should be taken explicitly rather than drifted into. Which surviving workarounds are being retired and by when. Which design gaps are accepted as they are and which become a change. And which measures graduate into business-as-usual reporting permanently, because the useful ones should not stop when the project does.
What ninety days cannot tell you
Two things, and it is worth saying them so that nobody over-reads the numbers.
It cannot tell you whether the design will hold at volume. A system that copes with a normal month has not been tested by a peak season, a year-end, or a statutory audit. The genuinely hard confirmation comes at the first close under pressure, which for most companies is somewhere between month four and month nine.
And it cannot tell you whether people believe it. Belief is visible in behaviour rather than in a survey: whether the operations meeting argues about the system's numbers or about a spreadsheet's, whether a manager asks for a report or builds one, whether new hires are trained on the system or on the workaround. Those are observations, not measures, and someone senior should be making them deliberately.
The period after go-live is also the period in which attention drains away fastest — the consultants demobilise, the sponsor returns to their day job, and the workarounds harden into habits while nobody is watching. That is the argument in why the most dangerous day of your project is not day one, and measurement is the only practical defence against it, because a number in a meeting keeps the attention that goodwill does not.
If the weekend that produced these numbers is still ahead of you, the sequence and the go criteria are worth writing first, and that is the cutover weekend. If it is behind you and the numbers are not being produced by anyone, an independent read of what your system is actually doing is a small piece of work with a short payback, and it is where an IT governance engagement usually starts.
