The biggest mistake is making this decision during the incident. Everyone is in panic mode, and whatever gets decided then is inconsistent and contested. This has to be settled long before anything breaks.
That means a predefined, tiered list of services, built top-down from the business - not a judgment call someone makes in the moment. Tier 0 is your most critical services: they directly touch the customer, and their failure carries serious business, media, or legal exposure. Tier 1, 2, and 3 step down from there. Because the tier is fixed in advance, it decides everything else automatically: tolerable downtime, who gets paged, what freezes, whether the postmortem happens today or next week. Nobody should be debating severity while the site is down.
For a Tier 0 incident it's all hands, and the cleanest move is a blanket moratorium on production changes - not just the affected service. That sounds heavy-handed, but it's not a work stoppage: development and testing continue, people keep building in their own environments, and merges resume the moment it lifts. What you're avoiding is compounding variables. If people keep shipping while you're working a critical fix, you multiply the pathways you have to investigate, and a new change may break something else - now you're debugging two problems at once. A temporary pause on pushing to production is far cheaper than that confusion.
The mitigation should already be built and rehearsed: divert traffic to a healthy replica, fail over, shed load. If you're inventing containment mid-incident, you've lost time you can't get back.
Communication follows the same principle. Customers should never discover an outage themselves - they are not your testers. You tell them early, on every channel, before you know root cause. Trust is built by being proactive, not by having a clean record.