Four iPhones running hotter than the rest
A team running a TikTok matrix hit this last summer.
Four iPhones on the rack were markedly hotter than the rest of the batch. Their first instinct was the script they had just changed, so they rolled back two versions. Still hot.
Two more days of investigation and they opened one up. The battery had swollen. All four came from the same purchase batch and had been in service eleven months.
The post-mortem found that the real problem was not being slow to notice. It was that the order was reversed: software suspected first, hardware last. One day spent pulling a single unit and feeling the battery would have saved the other two.
We already have an article on the long-term effects of running an Apple cluster control deployment, and that one covers making devices less likely to fail — brightness, thermals, rotation, cables. That is prevention.
This one covers the other half: it has already failed, now what. Prevention can be done gradually. The moment a device fails, you need an order of operations. And most people get that moment wrong — not because they do too little, but because they do it in the wrong sequence.
Three fault types, three diagnostic paths
Problems on a device rack look varied, but they sort into three kinds:
- Will not connect: the computer does not see it, the console shows it offline, or it works intermittently.
- Abnormal heat: noticeably hotter than its peers, or hot enough to reboot itself.
- Physical damage: a swollen battery, a bad charging port, a deformed screen or frame.
Why sort first? Because the handling logic differs completely. Most connection failures are software or contact problems and cost little. Abnormal heat requires separating load from hardware. And within physical damage, one item — swelling — is a safety issue with the highest priority.
Handle them jumbled together and the usual outcome is taking an entire group apart and re-pairing it to fix one connection problem, or discovering a device running absurdly hot, suspecting the script, spending two days on it, and finally finding a worn battery.
Will not connect: rule out the soft side first
Of the three, this is the most solvable, provided you do not reverse the order.
Recommended sequence, cheapest first:
- Swap in a cable you know works. Easiest step, most often skipped — the assumption is “the cable is new, it cannot be the problem”. But the cable is the most fragile link in the chain, and one that never moves can develop an internal break.
- Swap to a port you know works (another on the machine, or another socket on the hub). This separates “device problem” from “that port problem”.
- Try another computer. If it works elsewhere, the problem is on the original machine — driver, power, or USB controller.
- Check licence and environment status. An expired licence, developer mode not enabled, or the automation environment not started all present as “will not connect” and cannot be fixed from the device side.
- Only then suspect the device.
Reverse that order and you spend half a day dismantling the rack and re-pairing, to find it was the cable. Putting the most expensive diagnosis last is the entire point of the sequence.
One detail worth remembering: if the failure is intermittent, it is almost certainly a contact problem — cable, port, spring contact — not software. Software failures are consistent: working means it keeps working, broken means it never works.
Abnormal heat: find a reference first
Heat is the easiest thing to misjudge, because one device on its own gives you no idea whether that temperature is high.
Find a reference: same batch, same task. If only a few are markedly hotter, those few are the problem. If the whole batch is equally hot, the problem is the load, not the devices.
Once you have isolated the outliers, split further:
- Case hot, device otherwise normal. Usually a blocked thermal path — devices tiled too densely, the bottom vents covered by a stand, or dust inside. Easiest to fix: move it, leave a gap, clean once.
- Hot, plus spontaneous reboots, dropped frames or stutter. This needs a step further back. Rising internal resistance from an aged battery makes discharge generate markedly more heat. Check battery health; at a low enough level, plan a replacement.
- Only a few hot, and all bought in the same batch. The most concerning pattern, because it suggests a batch defect. If the cells or the mainboard thermal design share a flaw, they will show it the same way.
As for ordinary heat under normal load, that belongs to prevention — scheduling idle windows and improving physical cooling were covered there and are not repeated here.
Swollen battery: nothing to weigh
The earlier categories involve trade-offs. This one does not.
The moment you find a swollen battery, cut power and take it out of service.
The reason is not “performance may suffer”, it is safety. Swelling means irreversible chemical change inside the cell, and continued charging carries a fire risk. Cluster devices are precisely the ones plugged in around the clock, which multiplies that risk — especially during the hours when nobody is in the room.
Three actions:
- Cut power, not standby. Unplug the charging cable; power down if you can.
- Isolate it. Do not leave it pressed against other devices on the rack. Swelling pushes the case and screen outward, and continued pressure causes worse deformation.
- Deal with it soon. Do not leave it running tasks because “it still boots”. This is the one fault type where I would stop the business first and handle it.
How to catch it early? Visually, mostly — a device that rocks when laid flat on a desk. Every so often, pick each device up and look at the back for a bulge, particularly anything over a year old. Ten seconds of work that prevents something considerably worse.
Charging ports and screens: most common, easiest to defer
Neither stops a device immediately, which is exactly why both get deferred.
A bad charging port: clean it yourself first, then consider a repair shop. A fair share of “not charging” is dust, lint or oxidation in the connector, recoverable with compressed air and a gentle prod from something non-metallic. If it still fails after cleaning, the contacts are oxidised or the spring contact deformed, and that needs a shop.
Worth remembering: cluster devices stay plugged in for long stretches, so the connector wears far faster than on a daily-use phone. When one starts charging only sometimes, treat it as a device in the process of failing rather than an occasional glitch. Early handling is cheap.
Screens and frames: two main problems — clamp-style holders leaving marks, and edge pressure deforming the panel. Neither affects the script (scripts generally do not depend on the physical screen), but both hit resale value. If a device is visibly deformed, stop clamping it and change the mounting method.
Repair or replace: three lines
By this point you need a decision, not a technique. Three lines, any one of which means replace:
- The repair costs more than the same model used. The most direct test. Board-level repair frequently exceeds the price of a used replacement.
- The device holds data or configuration that is hard to migrate. An account binding, or parameters tuned individually. Here repair is often the better value.
- The fault affects batch scheduling. If the device sits in a group running synchronised tasks, an unstable member drags the group down. Replace rather than wait for the repair.
Conversely, what is worth repairing? A fault that is specific, single-point, and does not affect the group. A screen issue on a device that otherwise works, or a port that just needs cleaning. Those are cheap wins.
When to swap the whole batch
This is the most overlooked decision, because it triggers slowly.
Any one of three signals and a batch replacement deserves consideration:
- Failures start arriving in clusters within one batch — five or more in a month, say.
- The model is too old and the OS version no longer meets script requirements. Repair cannot fix being out of date.
- Failures concentrate in one production run or one age bracket. That is batch-level degradation, and the next one is already coming.
The cost of piecemeal repair is not only the repair bill, it is your time handling failures. Once they arrive in clusters, your cluster control operation has become an ongoing drain: you are always handling the next one instead of moving the business forward.
Two things to watch when swapping:
Prefer the same model. Different resolutions mean script coordinates need re-fitting. A uniform batch keeps maintenance an order of magnitude simpler.
Swap by group, not one at a time. A group carries the same tasks, so a group swap avoids tasks being unevenly distributed between old and new hardware, and lets you validate the script on new devices in a single pass.
Finally
Back to those four devices. With new batteries they ran another six months before retiring — because they were caught reasonably early, and because nobody kept poking at them afterwards.
Hardware failure handling is not about technical difficulty. It is about sequence and boundaries: rule out the cheapest possibility before touching the device; establish whether a fault is single-point or batch-level before deciding repair or replace; and stop immediately on swelling, with no trade-off.
Beyond that, most of it is avoidable through prevention — brightness, thermals, rotation, cables — and that article covers it in more detail than I do here: long-term running effects on controlled phones.
For what changes as the fleet grows, see how many phones one PC can actually control.
About EasyClick: A phone automation AI-agent platform covering Android no-root, iOS no-jailbreak (proxy / Bluetooth HID / OTG HID) and HarmonyOS Next, offering script development, Apple cluster control, local central control & mirroring, and cloud control systems. → Explore all products
Ready to build it for real?
Every approach in this article can be built on the EasyClick phone automation platform — full documentation, developer tools and cluster/cloud-control products, free to try.