By Luis Lopez, AI transportation consultant, CEO of Go Hub.io Holdings Corp and subsidiaries, and host of the Freight Guru Podcast
When a freight company tells me an AI tool “did not work,” the cause is usually not the tool. It is the data underneath. Freight data quality for AI matters because these systems learn patterns from your records and then act on them. If the records are inconsistent, the patterns are wrong, and so are the answers.
The good news is that you do not need perfect data. You need your most important fields to be consistent and trustworthy. Here are the ten I would check first, in rough order of how much damage bad data does.
The checklist
1. Customer and company names
The same customer entered five different ways looks like five customers. Abbreviations, typos, old names and duplicate accounts break reporting, pricing history and any analysis by customer. Pick one standard name per customer and merge the duplicates.
2. Addresses and location names
Pickup and delivery locations are often typed free-form. A warehouse might appear under several spellings, with and without suite numbers. Any tool that plans routes, estimates arrival or measures lane performance needs one clean record per physical location.
3. Reference numbers
Order numbers, container numbers, booking numbers, PO numbers and BOL numbers are how shipments get matched across systems. Check for stray spaces, wrong characters, and numbers placed in the wrong field. A tool cannot connect the dots if the identifier is not where it expects it.
4. Dates and times
Mixed date formats, missing time zones and appointment times typed into notes instead of fields cause a surprising number of errors. Use one format and give appointment, pickup and delivery times their own fields.
5. Equipment types and sizes
If one person enters a trailer type one way and another enters it differently, the data cannot be grouped. The same goes for container sizes and chassis types. Create a short, fixed list of allowed values and use it everywhere.
6. Commodity and freight descriptions
Vague descriptions such as “misc” or “freight all kinds” tell a system nothing. Commodity details drive pricing, handling and compliance. At a minimum, standardize how you describe what is moving and keep special handling requirements in a dedicated field.
7. Weights, quantities and dimensions
Mixed units, missing values and guesses entered as facts all corrupt rating and capacity planning. Decide whether you record pounds or kilograms, pallets or pieces, and enforce it. Mark estimates as estimates.
8. Rates, charges and accessorials
Charges buried in notes or typed under inconsistent names make it hard to see what a move really cost. Fuel, detention, storage, and other add-on charges should each have a consistent label. If an AI tool is going to help with quoting or margin analysis, it needs charges recorded the same way every time. See top 10 freight brokerage tasks for an AI copilot for where this matters in practice.
9. Status and event updates
Status fields tend to be the messiest because many people touch them and they are updated late. If “delivered” sometimes means arrived, sometimes means unloaded and sometimes means paperwork complete, any predictive tool will learn nonsense. Define each status clearly, and record when it actually happened, not just when someone typed it in. This ties directly to predictive ETA and freight visibility.
10. Driver, vehicle and carrier identifiers
Which driver, truck or outside carrier handled a move is key to performance, cost and compliance analysis. Inconsistent naming, reused unit numbers and missing assignments make it impossible to trace what happened. Keep identifiers unique and tie each move to the people and equipment involved.
How to clean without boiling the ocean
You do not have to fix everything at once. A practical approach looks like this:
- Choose the use case first. If you want AI for dispatch, prioritize locations, equipment and times. If you want it for billing, prioritize charges and reference numbers.
- Sample your data. Pull a few hundred recent records and look at them with your own eyes. You will see the patterns of mess quickly.
- Fix the source, not just the symptom. Use dropdown lists, required fields and validation so new bad data stops arriving.
- Assign an owner. Someone has to be responsible for each critical field, or it will drift again.
- Document your definitions. Write down what each status and field means, so everyone enters it the same way.
What cleaning cannot do
Be honest about limits. Cleaning fixes inconsistency, but it cannot create information that was never recorded. If nobody tracked actual stop times for the past two years, no amount of tidying will produce them. Our article on what AI cannot do in freight data quality covers where software helps clean data and where it only makes assumptions.
Questions to ask a vendor about your data
Once your fields are in better shape, put these questions to any vendor before you commit:
- Which of my fields does your product depend on most?
- What happens when a field is missing or wrong? Does it flag the problem or quietly guess?
- Can I see which records the system relied on for a given answer?
The guide to evaluating an AI vendor expands on these.
The bottom line
Before you spend money on an AI tool, spend a few weeks on these ten fields. You will get better results from any product you choose, you will be able to judge vendors more clearly, and you will find problems in your operation that have nothing to do with AI. Clean data is the cheapest improvement available to a freight company, and it pays off whether or not you ever buy the software.
For more on freight and technology, subscribe to the Freight Guru Podcast.
About the author: Luis Lopez is a Miami-based AI transportation consultant and logistics entrepreneur, the CEO of Go Hub.io Holdings Corp and subsidiaries, and host of the Freight Guru Podcast.


