1. Element Finding Is What Makes a Script Stable
When you write HarmonyOS automation scripts, the time rarely goes into business logic. It goes into making HarmonyOS scripts recognize the button that is on screen.
Hard-code the match and the script breaks the moment the interface changes. Loosen it too far and the script taps the wrong place. HarmonyOS Next offers a fair amount of capability here, and the usual problem is that people pick one method and force it onto every situation.
This article separates the four paths: image matching, color comparison, OCR recognition, and YOLO, and covers what each is good for, how to tune it, and the fallback order when a match fails.
2. What Each Path Is For
| Method | What it reads | Best for | Main cost |
|---|---|---|---|
| Image matching | Pixel features | Distinctive shapes, moving positions | Screenshot and match time |
| Color comparison | Color values | Single-color elements, fixed layout | Weak against style changes |
| OCR recognition | Text content | Text elements whose wording changes | Recognition time, model dependency |
| YOLO | Object category | Varying shapes and positions | Model file and training |
The four are not exclusive. Mature scripts mix them: image matching to confirm the page loaded, OCR to read the key text, and color comparison as a final check.
3. Image Matching
The most common technique in image recognition is image matching. It takes a small pre-cropped image as a template and slides it across the full screenshot, treating any position above the similarity threshold as a hit. HarmonyOS uses OpenCV for this, and the vendor cites a match rate above 95 percent.
Two things come first. Initialize OpenCV, and switch image storage to mat format, which measured memory down by half to eighty percent, CPU down twenty to thirty percent, and speed up one to two times. In HarmonyOS cluster control, where one script runs across dozens of devices, that gap becomes obvious.
Three practical notes on templates. Crop templates on the target device rather than reusing one model for another. Do not crop too small, because small templates match everywhere and a slightly loose threshold produces false hits. And keep changing text out of the template, leaving that job to OCR.
For the threshold, start from a sensible default and read the similarity the script actually reports, then tighten it until false hits stop. Different resolutions may need different values.
4. Color Comparison
Color comparison reads color values, which makes it far lighter than image matching.
Single-point comparison suits a fixed position whose state you want to check, such as a button that changes color when it becomes tappable. Multi-point comparison solves uniqueness: a single color may appear in many places, so binding several colors with relative offsets into one group narrows it down sharply.
There is also a reverse pattern: find by color first to narrow the region, then run image matching inside that region. That is much faster and more accurate than matching the whole screen.
Color comparison is brittle against change. A new accent color, a dark mode switch, or a different background image can break it. Treat it as a fast filter and a fallback, not as your only check.
5. OCR in Practice
OCR reads text, so it naturally handles elements whose wording changes. If the button label is rewritten, image matching needs a new template and OCR does not.
PP-OCRv6_small is built in and fine for upright interfaces. Initialization selects the model type, and the result comes back as an array where each entry carries the text, a confidence value, and a coordinate range you can tap.
Tune in this order. Padding first, because expanding a white border around text boxes fixes partial captures immediately. Then the detection box confidence threshold, where raising it reduces recall but improves precision. Then maxSideLen, capped at 640 for speed at the cost of very small text. Finally text direction detection and angle voting, which are only needed for rotated images.
One detail worth knowing: initialization failures expose a separate error message. Do not only check the boolean. Print the error and you will know immediately whether it is a model path problem or a parameter problem.
6. When YOLO Is Worth It
YOLO handles what the other three cannot: targets whose shape and position vary, but which you can describe as a category, such as detecting whether a particular icon appeared.
It adds a model file, and initialization and tuning are more involved than the other paths. Training follows the same approach as on Android. Unless you genuinely need it, do not reach for it first. The test is simple: if image matching, color comparison, and OCR together keep the flow stable, you do not need a model.
7. Debugging a Failed Match
Work through this order and you will usually find the cause.
Confirm the screenshot was captured. If capture returns nothing, everything downstream fails, so check this first.
Confirm the coordinate systems match. HarmonyOS taps and recognition both use screenshot coordinates, so if the screen size was set midway, verify both sides refer to the same thing.
Confirm OpenCV is initialized and the mat switch succeeded.
Then log the matched coordinates together with the screenshot. Many “element not found” reports turn out to be elements found at the wrong position, which is obvious the moment you look at the image.
One habit worth building: save the screenshot before the script exits on failure. Looking at the scene beats reading logs afterwards.
About EasyClick: A phone automation AI-agent platform covering Android no-root, iOS no-jailbreak (proxy / Bluetooth HID / OTG HID) and HarmonyOS Next, offering script development, Apple cluster control, local central control & mirroring, and cloud control systems. → Explore all products
Ready to build it for real?
Every approach in this article can be built on the EasyClick phone automation platform — full documentation, developer tools and cluster/cloud-control products, free to try.