← Blog/Operations

How to Evaluate an AI Agent for Freight Forwarding

Evaluating an AI agent for freight forwarding: the six parts a real platform has, the six questions to ask a vendor, and where the work should stop.

`} caption= /> )}

To evaluate an AI agent for freight forwarding, ignore the model underneath and check where the work stops. A real agent starts itself when a shipment email arrives, reads the documents, matches them against your existing records, produces the finished piece of work such as a created job or a drafted ISF, then hands it to an operator to approve. Software that returns a screen of extracted fields for someone to key is an extraction tool wearing an agent label, and the keying was the expensive part. The difference is measurable. At one US import freight forwarder running agents in production, the keying removed was worth about 17.5 hours a week across 25 jobs and 20 filings, with a person still reviewing and approving every record before it reached the TMS or a filing portal.

Every logistics vendor added the word “agent” to their product page over the last eighteen months. Some earned it. This is how to tell before you sign.

What makes a program an AI agent?

Four things, and a product missing any one of them is something else with a new label.

A trigger it did not ask you for. An email lands. A vessel departs. It is 6 a.m. The program starts itself, and nobody opens it, uploads to it, or clicks Run.

Context about your business. Your customers, your codes, your naming, the way your team writes a reference number. Without this a program can read a document but cannot tell you what to do about it.

Tools it can use. Read a PDF, query a database, log into a portal, create a record, send an email. A language model on its own produces text. An agent produces changes.

A place it stops. Well-built agents stop at a person. Badly built ones stop at whatever they were confident enough to do.

The same distinction as a comparison, because these four categories get sold interchangeably:

ToolStarts by itselfProducesBreaks when
ChatbotNoText you act onYou stop prompting it
OCR / extractionSometimesFields to key inNever; it just leaves the typing to you
Scripted macro / RPAYesA completed click pathA carrier changes a template
AI agentYesA finished record for approvalIt is unsure, at which point it should ask

How does an AI agent process a shipment, step by step?

Every agent worth the name runs the same six steps, and the value is concentrated in step three.

  1. Trigger. A packet lands in the ops mailbox.
  2. Read. Identify the documents, then pull the fields. Bill of lading, commercial invoice, packing list, worksheet.
  3. Resolve. Match what it read against what you already know. Which shipment is this, which of your customers is “SHANGHAI XX TRDG CO LTD,” which of your codes applies.
  4. Decide. Apply the rules. Notice what does not add up, and what is missing.
  5. Act. Produce the output: a created job, a drafted filing, a written reply.
  6. Hand off. Put it in front of a person with the source document attached.

Reading a bill of lading is close to solved and every vendor demos it. Knowing that the shipper on this bill is the same entity your team spells three different ways, on the shipment that arrived under a different house number, is the twenty minutes. Step six is where the liability sits, which is why it belongs to a person.

The six-step agent loop, from inbound email to approved record01 / 02Trigger and readThe packet arrives, the agent identifies the documents03 / 04Resolve and decideMatch to your job, your customer, your codes05ActA created job, a drafted filing, a written reply06A person approvesNothing is filed or written without this step

What is a serious AI agent platform built out of?

Six parts, and you will feel the absence of any one of them within a month of going live. This is the checklist to take into a demo.

  1. A versioned agent definition. What triggers it, what it reads, what rules it applies, what it may do, with a draft state and a publish step. Without versioning you cannot tell whether last Tuesday’s bad record came from the documents or from a change somebody made to the agent.
  2. Runs and steps you can open. One execution is a run; every action inside it is a step showing what it read, extracted, decided, and wrote. When an operator asks why it did that, the answer should be on a screen rather than in a support ticket.
  3. Knowledge that separates standing facts from per-shipment facts. Account facts stay true across every shipment for that customer: consignee entity, notify party, filing codes, reference formats. Shipment facts must be confirmed every time: this container, this weight, this vessel. An agent that treats the second like the first will reuse last month’s answer and be wrong in a way that looks right.
  4. A human decision step that pauses and resumes. Not a notification. The run stops, a person answers one specific question with the evidence in front of them, and the same run finishes. This is the difference between software that saves an operator time and software that adds a second inbox to their day.
  5. Integrations to the system you already run. API where one exists, browser automation driving the portal where one does not. Browser automation works and is the most fragile piece in any of these products, because a portal owner can move a button on a Tuesday. Ask what happens that week.
  6. A durable audit trail. Every run, field, source document, and human approval, kept and searchable. In customs work that is what you produce when someone asks why a filing said what it said.

Underneath all six sits the constraint that separates a compliance product from a demo. The agent prepares, a person commits. Nothing is filed with a government, sent to a customer, or written to the system of record without a human looking at it first. Late or inaccurate ISF filings carry CBP liquidated damages of up to $5,000 per violation, and that exposure lands on the filer of record, who is a person with a name.

Which freight forwarding jobs are worth giving an agent?

The ones that repeat, where the answer is already sitting in a document somebody emailed you, and where a reviewer can tell in seconds whether the output is right. That third condition does most of the filtering.

Back-office freight work has a cost curve that runs straight through headcount. Twice the shipments means twice the keying, which means another person, recruiting, training, and a slow ramp. Any forwarder past ten people knows the shape of it, and it is why peak season hurts. An agent bends that curve for the narrow class of task above and leaves everything else alone.

Tasks that fail the filter are the wrong place to start. Negotiating a rate with a carrier you have worked with for nine years fails it. Deciding whether to eat a demurrage charge to keep a customer fails it. Anything where the answer depends on a relationship or a judgment call is a bad candidate, and a vendor who says otherwise is selling ahead of what the technology does. In McKinsey’s State of AI research reported in March 2026, 23 percent of organizations were scaling an agent in at least one business function while no more than 10 percent had reached scale inside any single function. The gap between trying agents and getting value from them is, in freight, exactly the gap between extracted data and finished work.

Which agents can a forwarding operation actually run?

Name them after the job a person on your team does today. The industry has not settled on vocabulary, so the same thing sells as a job agent, an order agent, a lot agent, or a shipment intake agent depending on whose product page you are reading. Ignore the label and look at the work.

The job agent, split by lane. This is the workhorse, and it comes in variants because the paperwork changes shape with the trade lane.

Ocean import runs on the arrival packet: master and house bills, arrival notice, containers, and a charge set that has to match the port pair. TIO’s ocean import agent creates the job with parties, containers, and charges together. When a packet arrives holding two master bills it stops and emails a question instead of guessing, because one master bill is one job and merging them quietly is the error nobody catches until billing. It reads party identity off the labels on the page rather than off position, so a shifted layout does not move a field into the wrong box.

Air import runs on the MAWB and the flight manifest, on a shorter clock and with a flatter party structure than ocean, since no ISF exists on the air side. The detail that breaks naive air automation: a single master waybill can arrive split across several flights on different dates while the master record carries one ETA, so the arrival date has to be read off the manifest rather than the master. Across our air history 86.7 percent of masters carry exactly one house, which makes the remaining consolidations the cases worth flagging rather than processing straight through. Waybill numbers carry a modulus-7 check digit, so the agent validates the number arithmetically before trusting a read.

Domestic runs on the delivery order and the trucker’s paperwork, with a different charge standard and no customs clock. Same shape, tuned to its own documents, writing into the same system.

The compliance and filing agent. For US import that starts with the ISF 10+2. Most of its work is resolution rather than reading: which party on the packet is the seller and which is the manufacturer, matching a shipper spelled three ways, and picking the tariff code your team actually files for that product rather than the one a general classifier would reach for. TIO’s ISF agent reconciles the house bill against the worksheet, and when the two disagree about who the consolidator is, it stages the value from the carrier document and flags the conflict in the review email for the person submitting the filing. The output is a drafted filing your team reviews and submits, which took about 15 minutes by hand at the forwarder we measured.

The finance agent. Freight billing is high-volume and mechanical, which makes it good agent territory: splitting one job’s charges into separate invoices for the parties who actually owe them, reconciling a vendor invoice line by line against what was quoted, chasing an aging balance with a written summary. TIO’s finance agent does the invoice split today, an hour of careful clicking reduced to a review.

The audit agent. This one has no inbox trigger. It runs on a schedule, usually overnight, and re-checks work already done: filings against source documents, created jobs against the packets they came from, charges against the standard. Nobody would fund a person to do this, which is exactly why it is worth an agent. It is how you learn the system is drifting before a customer tells you.

And the ones past that. A pre-alert and arrival agent that watches for the vessel and sends the customer notice on your template. A track and trace agent that sweeps carrier portals and only speaks when something slips. A quoting agent that turns an inbound RFQ into a drafted rate. A carrier coordination agent that fans a drayage request out to several truckers, waits, and hands your operator one comparison table. A document chase agent that notices the packing list never arrived and asks for it. None of these are exotic. Each is the same six-step loop pointed at a different pile of email.

What should you ask a vendor before buying?

Evaluate where the work stops, not the model underneath it. Six questions, in the order that saves the most time.

  1. What starts it without a person? If the answer is “the operator uploads the file,” it is a tool rather than an agent. That may be fine. Price it accordingly.
  2. What is on the screen when it finishes? A completed record you approve, or a panel of extracted fields you retype. This one question separates most of the market.
  3. Does it write into the system we actually run? Ask for the name of your TMS out of their mouth, not the word “integrates.”
  4. What does it do when it is unsure? The good answer is that it asks a specific question and stops. The bad answer is a confidence score and a guess.
  5. Does “customs” mean drafted or checked? Several products describe customs work as filing readiness and field validation, meaning they confirm the required data is present and flag what is missing before the deadline. That is useful and it is not a drafted filing. Ask which one you are buying.
  6. What happens to a correction? If an operator fixes the same thing every week and nothing learns, you have bought an expensive intern.

Ask them on your own shipments. Any vendor confident in the product will run it on your real documents, and what comes back on your own messy paperwork tells you more than any accuracy figure on a slide.

Where does TIO fit?

TIO is an AI agent workforce for freight forwarders and 3PLs. The agents run the inbox-to-TMS layer: they read the shared mailbox your team already reads, do the prep across ocean import, air import, and domestic, and stage a finished record. Your team reviews and approves every one before it reaches the TMS or a filing portal, and corrections go back in as rules so the same fix does not get made twice. Nothing files, sends, or writes on its own.

If you want to see it on your own paperwork rather than a slide, book a demo and we will run the agents on your real documents.

Frequently asked questions

How do you evaluate an AI agent for freight forwarding?

Check where the work stops, not what model is underneath. A real agent starts itself when a shipment email arrives, reads the documents, matches them to the right job in your records, and produces a finished record for an operator to approve. A tool that returns a screen of extracted fields for someone to retype has left the expensive part of the job with your team. Ask what starts it without a person, what is on the screen when it finishes, whether it writes into the TMS you already run, and what it does when it is unsure. At one US import forwarder, the keying an agent removed was worth about 17.5 hours a week.

Do AI agents file ISF or customs entries on their own?

No, and no responsible product should. Liability for an ISF sits with the filer of record registered with CBP, and late or inaccurate filings carry liquidated damages of up to $5,000 per violation. Software cannot hold a filer code and cannot be sanctioned, so a person reviews and submits every filing. The agent does the reading, the party resolution, the classification, and the drafting. Several products in this category stop earlier still and provide readiness checks and field validation rather than a drafted filing, which is the single most useful thing to ask a vendor to clarify.

Does a freight forwarder need a different AI agent for ocean, air, and domestic?

Yes, because the paperwork changes shape with the lane. Ocean import runs on the arrival packet, master and house bills, containers, and a charge set tied to the port pair. Air import runs on the MAWB and the flight manifest, where one master waybill can arrive split across several flights on different dates, so the arrival date has to be read from the manifest rather than the master record. Domestic runs on the delivery order with no customs clock. A vendor selling one generic intake agent for every lane is telling you which lanes they have not built yet.

How is an AI agent different from OCR or document extraction?

Extraction gives you fields. An agent gives you a finished record inside the system you already run. The expensive part of the job is not reading a bill of lading, it is knowing which shipment it belongs to, which of your customers the shipper matches when three documents spell the name three different ways, and which of your codes applies. Extraction stops before all of that and hands the fields back for someone to type. That keying is what costs an ops team roughly 17.5 hours a week at 25 jobs and 20 filings.

Stop doing this by hand.

See an agent run it on your own shipments. Twenty minutes, no setup.

Book a demo