All Articles
Industry Insights

Where Freight Data Leaks: Five Gaps Between Email and TMS

Five specific places shipment data degrades between inbox and TMS, with the mechanism behind each leak and the defence that actually closes it.

Where Freight Data Leaks: Five Gaps Between Email and TMS

Every forwarder I've built pipelines for has the same story: the booking was right in the carrier's system, the packing list was right in the supplier's PDF, and somehow the shipment record in the TMS was wrong. Nobody typed it wrong on purpose. The data leaked out through a specific gap between systems. Here are the five I keep finding, with the mechanism behind each and the defence that closes it.

Leak 1: Retyping from PDFs

Mechanism. A packing list, commercial invoice or booking confirmation arrives as a PDF. Someone reads it and types the numbers into the TMS. There is no reason a human transcribing forty container numbers a day gets every one right, and no audit trail showing which typed value came from which document.

Defence. Extract structured fields directly from the document instead of retyping them, and keep the origin of every value:

{
  "field": "container_number",
  "value": "MSCU1234567",
  "source": "document",
  "document_id": "PL-2026-00931"
}

That source column matters more than the extraction itself. When something looks wrong later, you can trace it back to the exact document instead of asking "who typed this."

Run your own numbers: shipments per month × minutes spent re-keying and correcting a document field × your fully loaded ops rate per minute = the monthly cost of this leak alone. Most ops leads have never actually multiplied it out.

Leak 2: The packing-list-to-container mapping

Mechanism. A single packing list often covers several containers. When it gets processed as one flat table, quantities and marks blur across container boundaries. A carton that belongs to container two gets attached to container one, and nobody notices until a customs query or a claim forces someone to reconcile line by line.

Defence. Enforce one invariant at ingestion: one PDF, one container, one row. Every packing list is split at the point of extraction so each container gets its own row with its own line items, never merged with another container's cargo. If a packing list covers three containers, you get three rows, not one row with three containers jammed into a notes field.

You can check this on your own documents with the container validator before it ever reaches the TMS.

Leak 3: Party name mismatches

Mechanism. "ABC Trading Co" and "ABC Trading Co., Ltd." look like the same company to a human and look like two different strings to a matching system. Fuzzy matching tries to be helpful and links the shipment to the wrong party record, sometimes the wrong company entirely, because it scored "close enough."

Defence. Auto-link on exact name match only:

def link_party(candidate_name, party_records):
    exact = [p for p in party_records if p.name == candidate_name]
    if len(exact) == 1:
        return exact[0]
    return None  # send to manual review, do not guess

This looks stricter than it needs to be, until you've seen an invoice routed to the wrong consignee because a fuzzy matcher decided two similarly-named companies were the same party. A flagged review queue costs someone a minute. A misrouted invoice costs a lot more than a minute.

Leak 4: Overwritten carrier data

Mechanism. Manual entry happens first, because the booking desk needs to move. The carrier API confirmation arrives later, sometimes with different container numbers, ETAs or vessel names. Whichever system writes last wins, which means either the confirmed carrier data gets clobbered by an old manual entry, or a legitimate manual note gets silently erased because the API pushed a blank into that field.

Defence. A fixed source-priority hierarchy: liner API outranks manual entry, manual entry outranks document extraction. Documents and manual entries are only allowed to fill fields that are still blank. They never overwrite a value that came from a higher-priority source.

SourcePriorityCan overwriteCan enrich blanks
Liner API1 (highest)Manual, documentYes
Manual entry2DocumentYes
Document extraction3NothingYes, blanks only

The rule sounds obvious written down. It is the thing that's missing in most pipelines I've inherited, which is why carrier-confirmed data keeps quietly reverting to whatever a booking clerk typed three days earlier.

Leak 5: Fields that freeze at customs commencement

Mechanism. This one is specific to CargoWise: certain fields on a job freeze once customs commencement happens. If your upstream data was wrong before commencement and the correction only lands after, the field doesn't update. The TMS shows the old value forever on that job, and nobody notices because the job looks complete.

Defence. Validate and correct before commencement, not after. That means the four leaks above need to be closed upstream of the customs step, not patched afterwards, because after that point some fields are simply out of reach. If you're building or auditing a CargoWise integration and want the field-by-field detail on what locks and when, that's exactly the kind of thing covered in the eAdapter series.

Where this leaves you

None of these five leaks are exotic. They're mechanical: a human retypes a number, a flat table merges two containers, a fuzzy match guesses at a party, a late API write clobbers an early manual one, a status change locks a field. Each one has a specific point of failure, which means each one has a specific fix, not a vague "improve data quality" initiative.

If you want to see where your own packing lists are losing the container mapping, run one through the container validator and check the row count against your container count. That's a five-minute test that tells you whether leak two is happening in your pipeline right now.

Ready to automate your document processing?

Join freight forwarders saving hours every week with CargoMode.

Start for free