Every vendor of a container number recognition system quotes an accuracy figure. Almost none define how it was measured. That gap is where terminals lose money, because a number without a method cannot be compared between vendors or enforced against one. This article explains how to write an accuracy agreement that means something, and how to test it before you sign anything.
Consider two vendors. One quotes ninety eight percent, the other ninety five. On the face of it the choice is obvious, and on the face of it the choice is wrong.
The first measured character level accuracy in daylight on clean containers, excluding referrals to human operators. The second measured whole code accuracy across a full year including night traffic, wet weather and damaged units, counting referrals as failures. The second system is substantially better and the naive comparison picks the first.
Three definitions must be fixed before any figure becomes comparable. Whether the metric is character accuracy or whole code accuracy, since a single wrong character makes a code useless. Which conditions were sampled, covering weather, light and container condition. And whether referrals to a human count as failures or are excluded from the calculation. Docker Vision publishes more than ninety five percent accuracy for its recognition platform, and the useful conversation with any vendor starts by pinning down these three definitions rather than by comparing headline numbers.
Ask also over what period the quoted figure was achieved. A number from a controlled pilot on one lane over two weeks is a different claim from one drawn from a year of production traffic across several sites. Neither is dishonest, but only one tells you what to expect once your own gate is running in February with the light failing at four in the afternoon.

Container identification follows the ISO 6346 standard, which specifies an owner prefix, an equipment category identifier, a serial number and a check digit.
The check digit is the quiet hero of accuracy testing. It allows a system to validate its own read arithmetically rather than trusting the image alone, which means a well built recognition engine can often know it has misread before any human sees the result. Require any vendor to confirm they validate the check digit and to state precisely what happens when it fails. A system that reports a read as successful while the check digit does not reconcile is not doing the job you are paying for.
Prefix registration is administered by the Bureau International des Containers, whose explanation of the container identification number is the authoritative reference for the format. It is worth noting that related recognition tasks such as shipping document reading have no equivalent checksum, which is precisely why document accuracy is harder to guarantee than container code accuracy.
The standard also explains why certain misreads recur. Characters that are visually similar cause predictable confusion, and a container number recognition system that reports its confidence per character rather than only per code gives operators far better information when reviewing an exception. Ask whether per character confidence is available, because it materially affects how quickly a human can resolve a referral.
A test built from convenient samples will pass comfortably and then disappoint in production. The purpose of acceptance testing is to find problems, not to confirm a decision you have already made.
Build the sample from real gate traffic across a continuous period rather than from a curated set, and make sure it includes the difficult cases. Specify minimum proportions for night traffic, wet weather, containers with faded or damaged markings, and any container types unusual in your particular trade. If your terminal handles a lot of older equipment or a specific regional fleet, that must be represented.
Set the sample size large enough to be statistically meaningful rather than anecdotal, and run the test across enough calendar time to capture normal variation rather than one favourable week. State the ground truth process, meaning who determines what the code really was and how disagreements between the parties get resolved. Without that, a failed test becomes an argument rather than a result.
Decide in advance who runs the test and who holds the data. A test executed entirely by the vendor, on a sample the vendor selected, measured by a method the vendor defined, is a demonstration rather than an acceptance test. The terminal should at minimum control the sample selection, even where the vendor provides the tooling.

This distinction deserves its own section because it is the one most often missed and the one most easily exploited.
A recognition system has a choice whenever confidence is low. It can commit to a read and risk being wrong, or it can refer the case to a human operator. Referring more cases raises apparent accuracy on the reads it does commit to, while quietly moving work back to the people you were trying to free up.
The result is that a system reporting ninety nine percent accuracy while referring a third of traffic has not automated your gate, it has relocated the effort and added a screen. Always require both numbers, and set a ceiling on exception rate alongside the floor on accuracy. Then confirm how exceptions actually reach an operator, which the camera to TOS walkthrough illustrates, and make sure the workflow captures the correction rather than merely resolving the individual case.
Set the exception ceiling from your staffing reality rather than from an abstract target. Work out how many referrals an operator can genuinely resolve per hour alongside their other duties, multiply by available hours, and divide by expected traffic. That produces a ceiling you can actually staff, which is a far more useful number than one chosen because it sounds demanding.
The agreement should state the accuracy target, the measurement method, the sample definition, the exception rate ceiling, the review frequency and the remedy if targets are missed. Without a stated remedy the target is an aspiration rather than an obligation, and aspirations are not enforceable.
Include a retraining clause. Accuracy drifts as new shipping lines, new container types and unfamiliar markings appear in your gate stream, so the agreement should commit the vendor to periodic model updates rather than treating accuracy as a single acceptance event that happens once and is never revisited.
Set a review cadence, quarterly for the first year and then annually, with a defined process for raising a shortfall. Attach this section to the wider tender structure in the requirements checklist, and read it alongside the terminal automation buying guide so the accuracy clause fits the overall procurement rather than sitting apart from it. Industry publication Port Technology International is a useful source for tracking how recognition performance is reported across the sector.
Keep the remedy proportionate and practical. Service credits that are trivial will not change behaviour, and remedies so severe that no vendor would accept them simply push the risk back into the price. A staged remedy, escalating from a remediation plan through service credits to termination rights, tends to be both acceptable at signature and effective if ever invoked.
An accuracy figure for a container number recognition system is only as good as the method behind it. Define whole code accuracy, sample real traffic including night and adverse weather, separate accuracy from exception rate, validate the ISO 6346 check digit, and write both a remedy and a retraining obligation into the agreement. To discuss how recognition performance is measured and reported in practice, speak to the Docker Vision team.
Docker Vision publishes more than ninety five percent. The figure only means something alongside a defined method covering whole code accuracy, sampled conditions and how referrals are counted.
Character accuracy measures individual characters, whole code accuracy measures complete correct reads. Whole code is the meaningful metric, since one wrong character makes a code unusable.
A calculated digit within the container identification code that lets a system verify its own read arithmetically. Any recognition system should validate it and report failures.
Always. A system can reach a high accuracy figure by referring large volumes to human operators, which defeats the purpose of automating. Quote both numbers independently.
Large enough to be statistically meaningful and drawn from continuous real traffic rather than curated examples. It must include night, wet weather and damaged markings.
A well designed system routes the read to a human with the image attached, records the correction, and uses it for future retraining rather than silently guessing.
It can drift as new shipping lines and container types enter your traffic. Build a periodic retraining obligation into the agreement rather than treating accuracy as a one time test.
Modern deep learning models handle considerable degradation including rust, fading and partial occlusion. Severely damaged markings still route to exception handling, which is correct.
Define this in the agreement before testing begins. A stated ground truth and adjudication process prevents disagreements from stalling acceptance. See the cost variables.
28
Aug
Leave A Comment