The call at twenty to eight
Last winter a founder running home-goods exports called me at twenty to eight in the morning.
He had just opened his dashboard and found reach on three stores down to single digits. The day before it had been four figures. All three sat in the same device group behind the same egress.
He asked two questions: are we banned, and can I go and change the configuration now?
I told him not to touch anything yet. We walked the order instead — confirm whether the accounts could still log in, establish whether every device was affected or only some, and only then decide what to change. It turned out to be throttling rather than a ban. Four days to recover, and all three stores came back.
Had he acted on his first instinct, I doubt four days would have covered it.
This article is that order, written down. Two articles here already deal with prevention: one on anti-ban thinking for iOS cluster control, one on avoiding risk control in bulk phone automation. This one does not repeat them — it answers a single question: it has already happened, now what.
Prevention is a design problem. Handling is a sequence problem. The moment most people actually panic is when they open the console and find a batch of accounts whose reach has collapsed to near zero. What you need at that moment is an order of operations, not a mechanism.
One caveat: every platform judges differently, so what follows is judgement criteria and sequencing, not a fixed list for any specific platform. Fit this order to your own situation.
Step one: establish whether it is throttling or a ban
The two terms get used interchangeably, but the handling is not the same.
Throttling looks like this: the account still logs in, still posts, and no interface returns an error — nobody just sees it. In the numbers, reach, plays and impressions all collapse together. Some platforms give a vague notice; others say nothing at all.
A ban is login itself being refused. No interpretation needed.
Why sort this out first? Because throttling still leaves room to recover, whereas a ban means appeal or nothing. Throttling is also frequently the step before a ban. Treat it as “traffic has been poor lately” and leave it for a few days, and you may well arrive at the second stage.
So the first action on noticing abnormal data is to sample a few devices and accounts and confirm that login and posting still work. Five minutes, and it decides the direction of everything that follows.
Step two: stop everything, not just the affected devices
This is the least intuitive and most important step.
Instinct says rescue the failing batch first, so people immediately start editing the configuration on the affected devices, swapping networks, changing parameters, rerunning tasks. That reaction breaks two things at once.
First, it widens the fingerprint. Operating during handling is still operating. If the root cause has not been found, every step you take on the affected devices may deepen the problem.
Second, it destroys your control group. The most valuable thing you have is the devices that have not gone wrong — they are the only reference you have for working out where the cause sits. Keep running tasks while investigating, and by the time you look back the clean group is contaminated and attribution is no longer possible.
The correct order is therefore: stop everything, record, and only then touch anything.
Record at least four things:
- When the problem appeared, and when the numbers started to slide — these are often not the same moment
- Which devices and accounts are involved, and whether the scope is the whole batch or a cluster
- What tasks were running and with what configuration
- When the last change was made and what it was
That last item is frequently where the answer is, and the most commonly skipped — because everyone assumes “it was fine yesterday”.
Step three: isolate by device, not by account
Pick the wrong unit of isolation and every later judgement drifts.
Group by device. Operating behaviour, network egress and device environment are all tied to the hardware. Several accounts on one device usually behave alike; the same account moved to a different device is often fine.
Grouping by device quickly produces a table:
| Group | Behaviour | Reading |
|---|---|---|
| Group A devices | All affected | Device or network side |
| A few devices within a group | Affected, peers fine | Single-unit issue (cable, port, environment) |
| Entire batch | All affected | Most likely content or operating pattern, not hardware |
| Only devices on one account system | Affected | Account system or origin |
Why does grouping by account go wrong? Because accounts move. One running on group A today may be assigned to group B tomorrow. Group by account and you draw a net, not a boundary.
Step four: use controls to separate device, network and content
The basis for judgement is not speculation, it is controlled comparison. Three simple ones:
- Same content, posted on an unaffected device. Normal result → cause leans device or network. Same result → cause leans content or operating pattern.
- Are devices on the same network behaving identically? Collective failure points at the network egress; partial failure points at the devices themselves.
- Are recently reconfigured devices over-represented among the failures? If so, the answer is very probably in that change.
There is a precondition: you need devices you have not touched. That is the real reason step two insists on stopping everything. Stopping is not only damage control — it protects your one reliable reference.
It is also why I suggest keeping a small set of low-intensity accounts running day to day. They do little normally, but in a moment like this they are the only usable control group you have.
Step five: recover in batches, never all at once
Once handling is done, the recovery phase is where things most often go wrong again.
Releasing everything at once dumps every queued action from the handling window into one moment. The resulting density is more anomalous than whatever triggered the incident. Two hours after recovery, round two arrives.
The steady approach is three or four tranches:
- Start with the smallest-footprint, most stable accounts, and watch through a full business cycle — a whole day, for instance
- Confirm nothing retriggers, then bring on the next tranche at slightly larger scale
- Keep an observation window between tranches, and hold operating density below your normal pre-incident level rather than returning to full load immediately
- If a tranche fails again, stop scaling immediately and return to step three for fresh attribution
Recovering slowly costs nothing socially. Rushing back to full load usually means paying tuition with a second incident.
The three most common mistakes during recovery
One: operating while still investigating. Covered above, and the most frequent of the three. Urgency overrides process, and the trail is gone.
Two: replaying the original configuration unchanged. If your conclusion was that the configuration was at fault, recovery should carry the adjustment. Replaying it as it was simply retriggers the same conditions.
Three: nobody checks what changed recently. The worst thing during an incident is having the whole team on recovery. Someone should be tracing backwards in parallel: what changed on these devices in the past week — system, network, scripts, task scheduling. A great many root causes land on some change that looked trivial at the time.
Finally
Back to that call. He told me afterwards that what he was most grateful for was not acting on his first instinct — because he still had a dozen clean accounts, and those became the comparison group.
Incidents like this are rarely sudden. They are usually some accumulated change crossing a platform threshold. Which is why the end point of recovery is not “back to normal” but knowing what triggered it.
Do not skip the record. Notes taken during the incident are the only reference you will have next time: time and symptoms, devices and accounts involved, configuration and task state, actions taken and what changed afterwards.
Two further articles cover prevention and boundaries, and are worth reading first: the anti-ban guide for Apple cluster control and avoiding risk control in bulk phone automation. The compliance guide is also worth a look.
However you deploy, make sure the underlying use case is lawful and compliant. The technology is neutral; the boundary lies in the usage.
About EasyClick: A phone automation AI-agent platform covering Android no-root, iOS no-jailbreak (proxy / Bluetooth HID / OTG HID) and HarmonyOS Next, offering script development, Apple cluster control, local central control & mirroring, and cloud control systems. → Explore all products
Ready to build it for real?
Every approach in this article can be built on the EasyClick phone automation platform — full documentation, developer tools and cluster/cloud-control products, free to try.