Incident recovery
The restore step nobody puts in the runbook
Recovery plans are measured on how fast data returns to service. That single metric is why file hygiene falls out of the runbook, and why the document that caused the incident is frequently the one restored back into it. This piece sets out the mechanism, the incentive that hides it, and what a restore-hygiene step looks like when you add one.

Every recovery plan we have read is measured on how fast the data comes back. Recovery time objective, recovery point objective, hours of downtime, revenue lost per hour. Those are the numbers on the board slide, and they are the right numbers for most of what recovery has to do.
They are also the reason one step is missing from almost every runbook in the region.
The two dates
Start with the timeline, because the timeline is the whole mechanism.
You are compromised on one day. You find out on a later day, because that is what dwell time is: the interval between an intruder arriving and somebody noticing. Measure it in days or weeks rather than minutes. The exact figure varies by incident and matters far less than the shape of it.
Now ask which backup the recovery team reaches for. They reach for the most recent restore point believed to be clean. In practice that means the most recent restore point taken before the incident was detected, because detection is the only event with a timestamp everyone agrees on.
Those are two different dates. Everything created, received or modified in the gap between them is inside the backup, and the malicious document that started the incident was almost certainly created before anybody noticed anything. A backup is a faithful copy. It is faithful about the payload too.
The metric that deletes the step
Here is the part that gets less attention, and it is not a technology problem at all.
Recovery is scored on elapsed time. Every hour the estate is down is counted, often by people watching a revenue figure. Inspecting or reprocessing every restored document adds hours to precisely the number the recovery lead is being judged on, and it adds them at the worst possible moment, when the pressure to declare the incident over is at its highest.
Nobody writes "we skipped file hygiene" into a runbook. That is not how it happens. The step was never in the runbook, no metric ever asked for it, and the recovery completed on schedule. The incentive did the work quietly.
So the pattern repeats: servers rebuilt from clean images, credentials rotated, endpoints reimaged, network segmented properly this time. Then the fileshares, the mailboxes and the document repositories come back exactly as they were. Untouched. The only layer nobody reconstructs is the layer that carried the thing in.
Why re-scanning is the wrong instrument
The obvious answer is to scan the backup before restoring it, and it is worth being precise about why that is weaker than it sounds.
Your detection stack failed to identify this file once already, on the day it arrived. At restore time you are pointing the same class of engine at the same file, looking for the same catalogue of known-bad things. Signatures may have caught up if the campaign was noisy and widely reported. If it was not, or if the document was built for you specifically, they have not.
A sandbox has a related limit: it is only as good as it has been designed to be. A payload that waits for a user interaction, a specific locale, or a date that has not arrived yet behaves impeccably under detonation, and behaves impeccably for exactly as long as it needs to.
There is a quieter failure too. Re-scanning produces a verdict, and a verdict is not evidence. "Nothing found" is a statement about your detection coverage, not about the file.
What a restore-hygiene step looks like
The workable version treats a file coming out of backup the way you would treat a file arriving from a stranger, because during a recovery that is exactly its provenance.
Concretely, one step inserted between restore and return to service:
- Rebuild rather than inspect. Decompose each document into its ingredients, validate every one against the published specification for that format, then manufacture a brand-new file from that intermediate representation: the recipe, not the original bytes. The distinction matters more than it sounds. Copying across the parts an engine judges good is still a judgement about what looks safe, which is detection wearing different clothes, and the original file is still the one that arrives. A rebuild never passes the original through at all, so nothing survives by going unrecognised. This is what deterministic Content Disarm and Reconstruction does; Glasswall reports 100% of malicious files neutralised across 8.27 million tested, and its government research centre case study describes a recovery sequenced this way.
- Keep the record per file. For each document, what was removed, under which policy, at what time. This is the artefact your incident report needs and the artefact an assessor asks for, and it cannot be reconstructed later.
- Prioritise by blast radius, not alphabetically. Shared drives, mailboxes and anything an automated pipeline consumes go first. A finance share feeding a monthly macro-driven process is a higher-order problem than an archive nobody opens.
- Assume the estate is still islanded. Mid-recovery there is often no network to send files across, which is why desktop-resident processing matters here in a way it does not in steady state. Glasswall Meteor runs the engine locally, online, offline or air-gapped.
None of this is glamorous, and none of it is new engineering. It is one step and a decision to spend the hours.
The evidence half
The second reason to add the step is that in this region the obligation is already written down, and it is an evidence obligation as much as a control one.
Singapore's Cybersecurity Code of Practice for Critical Information Infrastructure spells it out in clause 6.1.4: logs kept for a minimum of twelve months after the event they record, protected against unauthorised modification and deletion, and "governed by a log retention policy to facilitate investigations into cybersecurity incidents". Retention whose stated purpose is the investigation, not the archive.
The MAS Technology Risk Management Guidelines set the parallel expectation for financial institutions: untrusted files arriving by mail, upload portal or third-party exchange are handled at the boundary rather than admitted on a detection verdict alone.
Read those together after an incident and the question an assessor arrives with is not "what did you block". It is "show me what happened to this document". A control that only records exceptions cannot answer that, because the files it allowed through generated no per-file record at all. A rebuild step produces one for every file it touches, including the ones that turned out to be perfectly fine, which is what makes the record usable as evidence rather than as an alert history.
The question worth asking
Recovery plans get rehearsed. Restore-hygiene steps get added after the second incident, which is a costly place to learn it.
If your signature goes on the recovery sign-off, the question for your team is not how fast the data came back. It is what did we put back, and how do we know.
For the full obligation picture in Singapore, including the Cybersecurity (Amendment) Act 2024 and the CCoP update announced in July 2026, see our Singapore compliance page.
Sources
- Glasswall: Content Disarm and Reconstruction, published testing across 8.27 million malicious files
- Glasswall case study: remediating a breach at a government research centre
- Glasswall Meteor: desktop CDR, online, offline or air-gapped
- CSA Singapore: Cybersecurity Codes of Practice for CII
- CSA Singapore: Cybersecurity Code of Practice for CII, Second Edition Revision One, clause 6.1.4 on log retention (PDF)
- MAS: Technology Risk Management Guidelines
See it against your own files
We will bring the engine, you bring the documents that matter. Contracted, deployed and supported in-region by Safeware.
Talk to us
