Finding one
The winning design is a decision stack, not a contest for one perfect rate signal
Ocean buying is a sequence of different decisions. Planning estimates lane demand and budget exposure. Market intelligence observes rates, capacity, congestion, connectivity, and disruption. Forecasting estimates how selected variables may evolve. Sourcing selects carriers, forwarders, allocation, rate mechanisms, and service terms. Booking secures space for specific cargo. Settlement verifies invoices and accessorials. A model, benchmark, or quote can support several steps, but none is authoritative for all of them.
AI earns value when it produces a calibrated range, explains which inputs moved that range, identifies where data are stale, and lets a buyer compare actions under uncertainty. Indexation earns value when both parties accept a precise reference, adjustment schedule, basis, collar, surcharge boundary, and fallback. Forwarder quotes earn value when they reveal executable terms and a provider can deliver capacity, documentation, exception management, and accountability. The procurement outcome comes from the combined policy, not the prettiest chart.
This distinction matters because maritime trade reacts to shocks through more than price. UN Trade and Development reported only 2.2 percent growth in seaborne trade volume during 2024 while ton-miles rose 5.9 percent as rerouting lengthened distances. The same physical cargo can therefore consume more vessel time and capacity even without equivalent volume growth. A model trained only on demand volume or a buyer focused only on a headline rate can miss the operational mechanism moving the market.
The practical recommendation is to build a transportation-and-logistics control tower around lane-level truth. Preserve forecast version, index observation, quote terms, awarded allocation, booking result, actual sailing, accessorials, invoice, and service outcome as separate evidence. That lineage makes it possible to learn whether an error came from prediction, procurement, capacity execution, routing, contract interpretation, or downstream operations.
Evidence: unctad-rmt-2025, wto-outlook-2026
Market object audit
Before comparing numbers, establish whether they describe the same corridor, box, and cost boundary
A container rate is not a natural constant. It describes an origin and destination geography, port or inland scope, direction, equipment size and type, cargo class, service and routing, shipment timing, validity, volume and commitment, payer terms, included surcharges, excluded terminal and inland charges, free-time terms, documentation, and commercial relationship. Two numerically identical rates can have materially different total cost and service exposure. A model comparison that ignores those dimensions can optimize the wrong target with great precision.
The Shanghai Containerized Freight Index illustrates scope discipline. The Shanghai Shipping Exchange describes SCFI route rates as Shanghai export spot ocean freight plus specified maritime surcharges for base ports, with defined units and CY-to-CY terms. It also lists charges that are not included, such as terminal handling at both ends and inland on-carriage. That methodology makes the index informative within its definition. It does not make the composite index an executable all-in quote for a shipper's exact origin, destination, commodity, service, equipment, or accessorial profile.
The Freightos Baltic container indices use another defined object. The Baltic Exchange states that FBX reflects weekly spot rates for forty-foot containers across twelve global trade lanes, while its guide documents benchmark administration and input processing. A shipper can legitimately use such a benchmark for analysis or a properly designed contract. It must still match the relevant lane and understand the published methodology, licensing, governance, update, correction, and cessation terms.
Build a rate ontology before building a model. Store base ocean, bunker or fuel, security, peak, low-sulphur, congestion, equipment, documentation, terminal, inland, customs, detention, demurrage, and other components separately when the commercial source permits it. Preserve currency, unit, effective interval, quote validity, offered capacity, service string, and source. Do not train on a field called total when its definition changes by supplier and month.
| Dimension | Question | Failure if ignored | Required record |
|---|---|---|---|
| Geography | Which origin, destination, base ports, inland points, and direction? | A global or port index is applied to another lane | Normalized corridor plus original wording |
| Equipment | 20-foot, 40-foot, high cube, reefer, special, or other? | TEU and FEU values or dry and special equipment are mixed | Container type, size, and conversion rule |
| Cargo | FAK, named commodity, dangerous, overweight, or controlled? | The benchmark or quote excludes the actual cargo | Commodity rule, weight, restrictions, and approvals |
| Rate boundary | Which surcharges, terminals, inland, documentation, and accessorials are included? | Base-rate saving becomes higher landed freight cost | Included, excluded, indexed, fixed, and pass-through components |
| Time | Observation, publication, booking, gate-in, sailing, or invoice date? | Future information leaks into a backtest or validity expires | All timestamps, time zone, validity, and revision state |
| Commitment | Spot, named account, service contract, MQC, allocation, or indicative? | A market indicator is mistaken for executable capacity | Volume, commitment, allocation, acceptance, and remedies |
| Service | Direct or transshipment, schedule, reliability, rollover, and exception handling? | Lowest rate wins while service failure creates other cost | Routing, service promise, performance, and escalation terms |
Evidence: sse-scfi-faq, baltic-fbx-guide, baltic-container-indices
Myth: a forecast is a buyable rate
Data finding: AI estimates a target variable; a commercial counterparty creates executable terms
An AI forecast can estimate next week's route index, the probability of a directional move, a quantile range, expected quote spread, or a classification such as stressed versus normal. Each target supports a different decision. None is automatically a rate a carrier or forwarder will honor. The buyer still needs a counterparty, contract or quote, validity, equipment, capacity, routing, terms, and booking acceptance. A forecast screenshot cannot be tendered as cargo.
The Federal Maritime Commission's service-contract overview provides a concrete U.S. regulatory example. It describes a service contract as an arrangement in which a shipper commits a minimum quantity over a fixed period and an ocean common carrier or agreement commits to a rate or rate schedule and a defined level of service. That exchange of commitment differs fundamentally from a model prediction. Other jurisdictions and commercial structures differ, so qualified professionals must determine the applicable rules and form.
Forwarder quotes also require role clarity. The FMC distinguishes an ocean freight forwarder, which arranges space and documentation on behalf of shippers, from a non-vessel-operating common carrier, which holds itself out as a common carrier, issues its own transport document, and acts as a shipper in relation to the vessel operator. A procurement dataset that labels every responding intermediary forwarder can obscure who assumes transportation responsibility and which terms are executable.
The controlled workflow labels the forecast as decision support, the index as a referenced observation, the quote as an offer subject to stated terms, the award as a sourcing decision, and the booking confirmation as an operational commitment. Each state has an owner and timestamp. That separation prevents an AI agent from auto-booking against an indicative number or a buyer from holding a provider to a prediction it never promised.
Evidence: fmc-service-contracts, fmc-oti
Myth: lower forecast error guarantees savings
Data finding: statistical accuracy and procurement value can move in opposite directions
Mean absolute error, root mean squared error, and percentage errors summarize prediction differences under particular denominators. A model can improve those scores by predicting ordinary weeks more closely while still missing the few shock periods that dominate budget exposure. It can also forecast an index accurately while the shipper's executable quote basis moves differently. Conversely, a modest forecast can support a valuable decision when it correctly signals a broad risk regime and the organization has a feasible response.
Procurement value requires a decision simulation. On every historical forecast date, define what the buyer knew, which contracts and capacity options were available, when action could take effect, how much volume was eligible, which transaction and switching costs applied, and what service constraints limited choice. Compare the policy using the model with a frozen baseline such as scheduled tendering, index-linked allocation, incumbent quote process, or a seasonal-naive forecast. A model cannot receive credit for actions the business could not execute.
Realized savings also needs a defensible counterfactual. Comparing an awarded rate with a later spot index can create false savings because the contract included capacity, service, different timing, or different components. Comparing one forwarder with the second-lowest quote can ignore routing and allocation. Report comparable price variance, total freight cost, capacity fulfillment, rollover, transit, accessorials, expedites, and procurement effort separately. Then state which outcome the forecast plausibly influenced.
A decision-weighted loss function can reflect consequence without hiding it. A team may assign higher loss to underpredicting a stressed regime when uncontracted launch inventory is exposed, but governance must prevent the weighting from becoming a machine for overbuying fixed capacity. Evaluate calibration and regret across scenarios. The best model improves a controlled policy on held-out periods; it is not merely the one that wins a statistical leaderboard.
Myth: an index removes volatility
Data finding: indexation changes how volatility is shared and measured
An index-linked agreement does not make the market stable. It defines how a contractual rate responds to a benchmark. Depending on cadence, lag, base period, multiplier, adjustment, floor, ceiling, trigger, and reset, the shipper can retain much of the movement, cap selected extremes, or exchange fixed-price renegotiation risk for transparent periodic change. The supplier can reduce the risk that a fixed rate diverges far from its buy or market basis. Both parties still carry volume, capacity, service, basis, and counterparty risk.
The index must be specified like a production dependency. Name the administrator, exact series and corridor, unit, publication time, subscription and usage rights, correction policy, holiday treatment, rounding, effective lag, missing-publication fallback, methodology-change review, and cessation replacement. State whether the formula uses level, change, moving average, median, or another transformation. A contract that says market index without this detail is an argument waiting for volatility.
Floors and ceilings need scenario tests. A narrow collar can behave like a fixed rate until the market leaves the band, then transfer movement abruptly. A wide collar can offer little budget protection. Adjustment lag can reduce noise but create mismatch during a rapid reversal. A quarterly average can smooth weekly peaks while failing to follow a supplier's short-term buy exposure. There is no universally fair formula, only a disclosed allocation of risk that fits or fails the lane.
Indexation can improve relationship durability when it reduces repeated winner-loser renegotiation, but that outcome still requires capacity and service terms. Define allocation, minimum quantity, forecast tolerance, booking lead time, rollover, blank-sailing treatment, equipment availability, and performance review outside the price formula. A benchmark can govern price movement; it cannot load a container or guarantee a sailing.
- Reference: exact route series, equipment, unit, source, license, and benchmark administrator.
- Transformation: level or change, base, multiplier, fixed adjustment, rounding, and currency.
- Timing: observation window, publication cadence, lag, effective period, and holiday treatment.
- Risk sharing: floor, ceiling, trigger, reopeners, extraordinary-event treatment, and termination.
- Continuity: correction, missing value, methodology change, cessation, and successor benchmark.
- Non-price terms: capacity, allocation, service, surcharge boundaries, performance, audit, and dispute.
Evidence: sse-scfi-faq, baltic-fbx-guide, xeneta-indexing-overview
Myth: the index is the lane
Data finding: basis risk is a measurable difference, not a footnote
Basis risk is the mismatch between the referenced benchmark and the shipper's actual commercial exposure. Geography creates basis when an index covers base ports while cargo moves from an inland origin to a secondary destination. Product creates basis when the benchmark reflects general dry freight but the shipper needs reefer, special equipment, named commodities, or overweight handling. Commercial basis arises when the index represents spot rates while the agreement includes annual volume, capacity, or service commitments.
Rate composition creates another basis. The SCFI methodology names included maritime surcharges and excluded charges. FBX has its own published benchmark definition. A shipper should not mix these objects casually or assume any global composite tracks its landed lane cost. Create a component crosswalk between benchmark and historical invoices or qualified quotes. Items outside the benchmark need fixed, indexed, capped, pass-through, or audit rules of their own.
Measure historical basis as actual comparable lane cost minus the contract's synthetic indexed value, using only aligned dates and components. Report mean, median, volatility, tail, trend, and performance during disruption. Stable average basis can still hide unacceptable extremes. Test whether basis correlates with secondary-port congestion, inland capacity, equipment imbalance, carrier mix, season, or surcharge changes. If the network changes, historical calibration may no longer apply.
Govern basis after launch. Set a review band and minimum sample before reopening the fixed adjustment. Prevent opportunistic resets based on one adverse month. Retain quote, invoice, index, mapping, and exclusion evidence so both parties can reproduce the calculation. If methodology changes materially, apply the contract review process rather than silently treating the discontinuity as market movement.
| Basis source | Diagnostic | Contract control | Monitoring signal |
|---|---|---|---|
| Port and inland | Compare base-port benchmark with port-pair and door components | Separate inland or document a geographic adjustment | Residual by origin and destination |
| Equipment and cargo | Slice dry, reefer, special, weight, and restrictions | Use a qualified series or exclude affected cargo | Quote spread by equipment and cargo |
| Spot versus committed | Compare benchmark tenor with award and capacity terms | Fixed service adjustment and explicit commitments | Basis by lead time and MQC performance |
| Surcharge boundary | Map each charge to included, fixed, indexed, or pass-through | Named component schedule and evidence requirement | Invoice variance by charge code |
| Timing | Test publication, averaging, and effective lags | Defined observation window and reset cadence | Basis around rapid rise and reversal |
| Methodology | Track corrections, constituents, weights, and definition | Change and successor clause | Structural break after method change |
| Service quality | Compare price with routing, reliability, rollover, and capacity | Separate performance terms from indexed price | Total cost and service by supplier |
Evidence: sse-scfi-faq, baltic-fbx-guide
Myth: a forwarder quote is merely a spot observation
Data finding: an executable quote packages role, validity, capacity, route, and exclusions
A forwarder quote should be normalized, not stripped of meaning. Record whether the provider acts as agent, NVOCC, or another role; who issues the transport document; which carrier or service is offered; whether space is firm or subject to confirmation; which sailing window and route apply; how long the offer remains valid; which cargo and equipment assumptions apply; and how exceptions are handled. A low headline line item with missing capacity or broad exclusions is not equivalent to a higher qualified offer.
For U.S. trades, the FMC OTI guidance explains distinct obligations and commercial roles for ocean freight forwarders and NVOCCs, including tariff publication for NVOCCs and possible negotiated arrangements. The point is not to provide legal advice or prescribe a form. It is to show why the procurement system must preserve counterparty role and governing terms rather than reduce every response to one rate field. Verify current status and requirements directly for the transaction.
Quote comparison should use a bid schema. Normalize origin and destination scope, equipment, commodity, weight, free time, transit, transshipment, sailing, validity, base ocean, surcharges, terminal, inland, documentation, customs-related services, detention and demurrage terms, insurance, taxes, and exclusions. Flag rather than guess missing cells. Let providers confirm the normalized comparison before award so a parsing error does not become a commercial dispute.
A quote process also reveals capacity and service intelligence that an index may not. Declines, short validity, limited allocations, changed routings, or rising exception language can be early signals, but they are not clean labels unless reasons are captured. AI can extract terms and detect anomalies, while a buyer validates material items and negotiates. The executable offer, acceptance, and booking remain controlled commercial records.
Evidence: fmc-oti, fmc-service-contracts
Myth: a global composite fits every portfolio
Data finding: aggregation is useful for context and dangerous for lane execution
A global index answers a global index question. It can summarize broad direction, support executive context, or become a feature when its relationship to a lane is empirically stable. It does not automatically represent a particular port pair, direction, equipment, season, service, or shipper segment. Opposing lanes can move differently, and trade-lane weights can cause a composite to rise while a smaller lane falls. Procurement action should use the most specific qualified exposure available.
Connectivity signals follow the same rule. UNCTAD's revised Liner Shipping Connectivity Index measures an economy's integration in the liner network using six components: ship calls, deployed capacity, services, companies, largest ship, and direct connections. It is valuable structural context. It is not a freight quote, schedule-reliability promise, or weekly price predictor. A connected country can still experience a port-specific disruption or shipper-specific equipment shortage.
IMF PortWatch provides open indicators based on satellite vessel data to monitor and simulate maritime disruption. IMF research describes port calls, AIS-derived shipment-volume estimates, and transit metrics while explaining limitations such as reception gaps, revisions, mode differences, and timing mismatches with customs statistics. These signals can enrich a model, but using revised PortWatch values in a historical backtest before their actual publication date would leak the future.
Build a hierarchy: portfolio and macro context at the top, corridor and direction below, port pair and inland scope below that, then supplier, service, equipment, and cargo. Features can flow downward only after a documented lag and relevance test. Decisions and evaluation should roll upward from the lane exposures that were actually buyable. This prevents a strong global relationship from masking a weak local basis.
Evidence: unctad-lsci-2026, imf-portwatch, imf-portwatch-research, baltic-container-indices
Myth: sophisticated AI always beats a transparent baseline
Data finding: published model rankings vary with route, sample, horizon, and target
Freight forecasting research does not support one permanent algorithm winner. Munim and Schramm compared ARIMA, VAR or VEC, and artificial neural-network models on CCFI routes. Their reported ranking changed between training and test samples and had route-level exceptions. A 2025 comparative study evaluated decision tree, random forest, Prophet, and LSTM approaches and reported strong decision-tree performance in its setting. Different data, periods, splits, features, and metrics can produce different winners.
A 2026 systematic literature review describes growing machine-learning use in freight-rate forecasting amid volatility and catalogs a wide range of methods and markets. That breadth is evidence of an active field, not a standard model that every shipper should deploy. Published studies can establish plausibility and design ideas. They cannot validate a private model on another corridor, contract basis, information set, horizon, or procurement policy.
Every production benchmark should include naive and operational baselines. Candidates include last observation, seasonal naive, moving average, simple exponential smoothing, scheduled market reset, and the current procurement analyst process. Use the same vintage data and forecast dates for every method. If a complex model cannot outperform a simple baseline after costs and instability on held-out periods, keep the baseline. Complexity is a liability unless it buys durable decision value.
Model selection should be conditional. A linear model may be easier to maintain in a stable regime; tree or neural approaches may capture nonlinear relationships when enough reliable data exist. An ensemble can reduce some variance but make lineage and change control harder. The policy can route to different models by horizon or lane, or abstain outside validated conditions. Procurement needs a reliable forecast service, not loyalty to an algorithm brand.
Evidence: munim-schramm-2021, kim-cha-jeon-2025, ml-freight-review-2026
Three-method scorecard
Forecast, index formula, and quote process answer different procurement questions
The comparison becomes productive when each method is scored against the question it can answer. AI forecasting is strongest for probabilistic market scenarios, anomaly detection, and timing support. Index-linked buying is strongest for a transparent and repeatable price-adjustment mechanism when basis is acceptable. Forwarder quoting is strongest for executable offers, capacity, routing, commercial terms, and operational accountability. A mixed portfolio can use a forecast to define risk posture, an index to govern part of price, and quotes to cover execution.
None is free of human work. Forecasts need target definition, data rights, validation, monitoring, and interpretation. Index contracts need benchmark due diligence, formula negotiation, audit, basis monitoring, and continuity rules. Quote processes need scope normalization, supplier qualification, bid analysis, negotiation, award, booking, and invoice control. Include these costs when comparing cycle time or return. A process that moves work from procurement into data operations has not eliminated the work.
Complete this scorecard per lane and season. A method that fits a high-volume benchmark-aligned dry corridor may fail a low-volume reefer lane with inland complexity. Preserve disqualifying constraints before weighted scores; an attractive price model cannot compensate for inability to carry the cargo or provide capacity. Portfolio standardization should govern evidence and decisions while permitting lane-specific buying mechanisms.
| Decision dimension | AI rate forecast | Index-linked buying | Forwarder quotes |
|---|---|---|---|
| Primary output | Future distribution, direction probability, scenario, or risk regime | Contract adjustment derived from a named benchmark | Executable offer subject to validity, capacity, routing, and terms |
| Best use | Budget risk, timing, allocation scenarios, and warning | Transparent periodic reset and shared market movement | Lane sourcing, tactical coverage, and execution |
| Main dependency | Vintage rate, demand, capacity, network, disruption, and event data | Reliable licensed representative benchmark and precise formula | Complete bid scope, qualified counterparties, and comparable responses |
| Main error | Forecast miss, miscalibration, leakage, drift, or wrong target | Basis mismatch, ambiguity, lag, method change, or cessation | Noncomparable scope, expired validity, unavailable capacity, or exclusion |
| Capacity access | No capacity by itself | Only what the associated contract provides | Can be offered and confirmed subject to terms |
| Price certainty | None; the output is an estimate | Formula certainty with a variable resulting level | Offer certainty during validity after acceptance conditions |
| Governance owner | Forecast product owner plus procurement risk owner | Contract, benchmark, finance, and procurement owners | Category buyer, supplier owner, operations, and invoice control |
| Success measure | Calibration, interval coverage, error, stability, and regret | Reproducibility, acceptable basis, durability, and service | Comparable total cost, capacity, service, exceptions, and invoices |
| Fallback | Baseline forecast, wider interval, abstention, and manual scenarios | Last value, successor benchmark, review, or termination clause | Alternate qualified supplier or contract allocation |
Manual procurement baseline
Benchmark AI against the quote-to-booking process the organization actually runs
A credible baseline begins with lane demand. Procurement receives volume forecasts, validates equipment and cargo constraints, decides which lanes enter a tender or quote event, builds a bid sheet, invites qualified providers, answers clarifications, normalizes responses, evaluates price and service, negotiates, awards allocation, and transfers terms into contract and booking systems. Operations then tests whether awarded capacity and routing materialize. Finance compares invoices with applicable terms. Forecasting is only one input in that chain.
Map labor and delay at every step. Buyers may spend little time interpreting market direction but many hours cleaning port names, aligning surcharges, chasing missing free time, or reconciling amendments. An AI rate forecaster cannot claim those hours unless it actually changes them. Conversely, automated quote extraction and normalization can create value even if rate prediction remains unchanged. Separate forecasting, document intelligence, sourcing workflow, and execution benefits.
The human baseline is not perfect. Buyers can anchor on an incumbent, overreact to one headline index, conflate rate and capacity, delay decisions, or compare inconsistent scope. They also know product launches, supplier relationships, network changes, cash constraints, service failures, and negotiation context missing from data. A fair test gives human and model processes the same vintage information, preserves independent judgments on a sample, and adjudicates why decisions differ.
Baseline metrics include tender cycle time, comparable responses per lane, quote validity at award, allocation awarded, booking acceptance, rollover, actual transit, exception rate, invoice variance, accessorials, expedite spend, and total cost per completed shipment. Add buyer and specialist minutes. A forecast experiment is incomplete if it improves an index metric while commercial outcomes remain unobserved. This operational connection belongs inside a broader supply-chain software design.
Evidence: fmc-oti, fmc-service-contracts
Data and leakage audit
The most impressive backtest may contain information the historical buyer could not know
Define the target first. A route-index level differs from a shipper's all-in executable quote, index change, quote spread, capacity acceptance, or landed freight cost. Training across these labels without provenance creates an incoherent target. Choose horizon and decision cadence: next week for tactical coverage, next month for an index reset, next quarter for sourcing, or annual scenario bands for budgeting. The information set and error measure change with each horizon.
Market features can include lagged route indices, rate components, quote distributions, bookings, trade statistics, capacity deployment, services, port congestion, vessel transit, equipment signals, fuel, currency, holidays, and declared disruptions. Every feature needs event time, publication time, revision time, geography, unit, license, and permitted use. A current dashboard export that overwrites history is not a valid training source. A data-management foundation must reconstruct what was actually visible at decision time.
Time leakage occurs when revised trade statistics, retrospectively labeled events, completed voyage data, final invoices, or backfilled index histories enter a row before their original publication. Quote leakage occurs when accepted quotes, awards, final allocations, or invoices enter features for a forecast meant to guide that same sourcing event. Related amendments split across train and test can let a model memorize supplier-lane templates rather than learn future behavior.
Preprocessing can leak too. Normalizing with full-period minima and maxima, selecting features on the entire dataset, imputing with later observations, or tuning thresholds on the final test contaminates evidence. Fit every transform within each training fold, use rolling or expanding windows, and lock a final untouched period. An independent reviewer should reproduce several forecast cutoffs from raw source records before the model can influence procurement.
Internal data create both opportunity and selection bias. A quote exists because a buyer requested it; awarded lanes receive more follow-up; final invoices appear only for moved cargo. Preserve requested, offered, declined, expired, awarded, booked, canceled, sailed, and invoiced states. Do not treat absent quotes as high rates or awarded rates as a random sample of the market. Selection is part of the commercial process and must be modeled or disclosed.
Evidence: imf-portwatch-research, nist-ai-rmf
Uncertainty and backtesting
A point forecast hides the procurement question inside false precision
Ocean markets combine ordinary variation with regime changes, rerouting, port disruption, policy shifts, and network responses. A single number cannot express whether the model sees a narrow stable distribution or a broad asymmetric one. Produce quantiles or prediction intervals at each horizon and test empirical coverage. If an eighty-percent interval covers only half of realized outcomes, it is not an eighty-percent operational interval regardless of its mathematical label.
Uncertainty should widen with horizon and data deterioration. Missing index observations, route changes, low quote density, unfamiliar events, or inputs outside training range can trigger wider bands or abstention. Procurement needs an actionable statement: expected range, tail scenarios, drivers, blind spots, and how much eligible volume is exposed. A confident median without that context encourages false certainty exactly when disruption makes the forecast least reliable.
Use rolling-origin evaluation. At each historical date, train only on then-available data, generate required horizons, and advance the origin. Report mean absolute or scaled error, root mean squared error where tail sensitivity matters, directional and regime performance, empirical interval coverage, width, interval score, and tail miss severity. Slice by corridor, direction, season, regime, horizon, equipment, and target. Compare naive, seasonal, statistical, machine-learning, and current analyst baselines.
Then simulate a frozen procurement policy. Measure normalized comparable cost, capacity shortfall, service consequence, quote or contract transaction cost, switching, unallocated volume, and regret relative to the best action that was actually feasible at that time. Perfect hindsight is an upper bound, not a fair baseline. The primary comparison is the current policy and simpler governed alternatives. Statistical significance alone cannot justify licenses, integration, buyer review, and model operations.
Peer-reviewed interval-forecasting research shows that uncertainty bands are an active objective in shipping-rate modeling, while comparative research shows model rankings vary. Those studies support the evaluation approach, not a universal production threshold. The organization's own decision clock, data vintage, basis, and capacity options determine whether a forecast is useful.
| Measure | Question answered | Required slice | Misleading shortcut |
|---|---|---|---|
| Point or scaled error | How far is the estimate from the defined target? | Lane, horizon, regime, and target basis | One aggregate score across dissimilar routes |
| Direction or regime | Does the model distinguish material rising, falling, or stressed states? | Magnitude band and decision lead time | Counting tiny moves as useful wins |
| Interval coverage and width | Do ranges cover outcomes without becoming uselessly wide? | Horizon, regime, lane, and data quality | Publishing uncalibrated confidence bands |
| Tail miss severity | How badly does the model fail in disruption? | Historical and synthetic stress sets | Optimizing normal weeks and hiding shocks |
| Decision regret | How did a frozen feasible buying policy perform? | Eligible volume, timing, transaction cost, and constraints | Using perfect hindsight as baseline |
| Capacity and service | Did the recommendation secure usable space and delivery? | Allocation, booking, rollover, transit, and exception | Crediting a low rate that was not executable |
| Stability and drift | Does performance survive market and data change? | Rolling residual, feature, and basis diagnostics | One launch backtest with no monitoring |
| Total decision cost | Does value exceed data, workflow, and governance cost? | Lane contribution and portfolio overhead | Ignoring licenses and buyer review |
Evidence: interval-forecasting-2026, munim-schramm-2021, ml-freight-review-2026, nist-ai-rmf
Illustrative procurement lab
A normalized portfolio shows why the forecast should change risk posture, not impersonate a price
Assume a fictional consumer-goods importer expects 12,000 FEU over the coming year on one qualified corridor. This is an illustrative planning volume, not client data. The team assigns 50 percent to a corridor-specific index-linked agreement, 35 percent to fixed or periodically quoted service allocations, and 15 percent to a tactical reserve. Those shares express one hypothetical risk policy, not a universal recommendation. Every contract retains its actual capacity, service, and surcharge terms.
For analysis, define the current qualified corridor benchmark as 100 normalized units. This is not a market price, currency amount, or published index value. The AI produces an illustrative eight-week median change of plus eight percent with an eighty-percent interval from minus six to plus twenty-seven percent. The interval crosses zero and is asymmetric. Translating only the median into buy now would discard both favorable downside and material upside uncertainty.
On the indexed half of volume, a ten-percent benchmark increase creates approximately five percent first-order indexed base-rate exposure at portfolio level before lag, collar, fixed adjustment, basis, and non-indexed charges. The fixed or quote-covered 35 percent has different exposure: price can be fixed during validity, while capacity, allocation, routing, and supplier performance remain. The 15-percent reserve preserves flexibility while accepting greater spot and capacity uncertainty. These mechanics require no invented dollar rate.
The buyer stress-tests feasible actions. One leaves allocation unchanged. Another advances part of qualified fixed coverage for launch-critical inventory. A third increases reserve because demand itself is uncertain. The decision record shows downside regret if the benchmark falls, upside exposure if it rises, capacity value, cancellation or volume risk, and service consequence. A procurement owner selects within policy and records a reason; the model does not execute a booking.
After eight weeks, the team evaluates the forecast against the exact index target and interval, the decision against feasible alternatives, the indexed calculation against contract terms, and quoted allocations against booking and service. It does not claim savings by choosing whichever later index observation makes the award look best. The example is intentionally normalized and contains no invented market price.
| Portfolio sleeve | Illustrative share | Price mechanism | Main non-price risk | Forecast role |
|---|---|---|---|---|
| Index-linked core | 50% or 6,000 FEU | Named corridor benchmark plus formula | Basis, capacity, service, lag, and collar | Stress reset exposure |
| Fixed or periodic quote | 35% or 4,200 FEU | Accepted offer or service-contract rate | Validity, volume, route, allocation, and performance | Timing and quote-spread context |
| Tactical reserve | 15% or 1,800 FEU | Future qualified quote or approved alternative | Spot volatility, capacity, and buyer workload | Tail risk and trigger |
| First-order sensitivity | 50% indexed share | A 10% benchmark move implies about 5% indexed base-rate exposure | Formula and basis change the actual result | Expose sensitivity, not guaranteed cost |
| Illustrative forecast | Eight-week horizon | Median +8%; 80% interval -6% to +27% | Calibration or regime can fail | Inform scenarios; no automatic booking |
Capacity, service, and failure investigation
A lower rate is not realized value when cargo rolls, reroutes, or arrives too late
Procurement must pair price with capacity and service evidence. Record awarded allocation, booking request, acceptance, rejection, equipment confirmation, rollover, blank sailing, routing change, actual departure, transshipment, arrival, and exception handling. Separate events a provider controlled from systemwide disruption and shipper-caused change. A forecast can anticipate stress, but capacity value comes from commercial terms and execution.
UNCTAD's 2025 review describes chronic disruption, rerouting, longer distance, delay, cost, and emissions. Those mechanisms can alter schedule and vessel utilization without mapping one-to-one into a weekly rate. A buyer may rationally pay more for a direct service, equipment access, or reliable allocation when launch consequences are high. The sourcing scorecard should expose that trade rather than let the cheapest normalized line item win automatically.
AI failures include wrong target, mixed scope, future leakage, uncalibrated intervals, low-density overconfidence, regime shift, and feature outage. Index failures include wrong corridor, unit, formula, lag, rounding, correction, methodology change, license breach, or cessation. Quote failures include exclusions, expired validity, wrong equipment, unavailable capacity, ambiguous provider role, parsing error, routing substitution, and amendments absent from booking or invoice systems.
Interaction failures can amplify each other. A model may train on synthetic index rates and present them as market observations. A quote parser may impute missing terms from the forecast, then the forecast later trains on those imputed quotes. Procurement may award more volume because a model predicts a rise, then interpret resulting quote scarcity as confirmation. Preserve observed, derived, forecast, offered, contracted, booked, and invoiced states separately.
Demurrage, detention, inland, terminal, and exception costs also sit outside many headline ocean comparisons. Applicable rules and terms must be checked directly. Treat total delivered shipment cost, capacity fulfillment, transit, exception work, and invoice accuracy as separate outcomes. Do not credit the forecasting model with avoided disruption unless its recommendation, executed action, and counterfactual are demonstrable.
Evidence: unctad-rmt-2025, fmc-demurrage-update, nist-ai-rmf
Human oversight, security, and governance
Give buyers decision rights while protecting commercially sensitive rate and volume data
Human in the loop is not a control description. Define who owns forecast use, benchmark selection, formula approval, lane award, capacity allocation, contract acceptance, booking escalation, and invoice dispute. A data scientist can validate prediction without authority to commit freight. A buyer can negotiate without authority to expose sensitive quote data to an unapproved model. Segregate roles according to consequence and preserve an accountable commercial approver.
Review interfaces should reduce anchoring. On sampled decisions, show market facts, demand, capacity constraints, and uncertainty before the recommendation. Show the baseline, recent errors, interval history, out-of-range warnings, and feasible alternatives. Capture a structured reason when the buyer changes timing or allocation: product urgency, supplier capacity, demand uncertainty, service failure, contract boundary, model concern, or another approved category.
Ocean intelligence can contain confidential rates, minimum quantities, named accounts, factories, product launches, margins, seasonal demand, invoices, and service weaknesses. Inventory fields, contractual restrictions, jurisdictions, providers, training use, retention, export, and deletion. Apply least privilege by entity, region, lane, role, and event. Separate data-source credentials from procurement execution, encrypt data, manage secrets outside prompts, and monitor bulk export.
External documents are untrusted inputs. A forged carrier notice or malicious attachment can poison an event feature or redirect a generative extraction agent. Treat documents as data, never policy instructions. Allowlist important sources, retain provenance, scan files, validate typed outputs, restrict tools, and require buyer confirmation for material quote and contract fields. A model must not email a supplier, accept an offer, or alter allocation because embedded text requested it.
NIST organizes voluntary AI risk work through Govern, Map, Measure, and Manage. Translate those functions into an owner, intended and prohibited uses, model and data inventory, impact mapping, validation, security review, release approval, monitoring, incident response, rollback, and retirement. A governed AI engineering practice makes each recommendation reproducible without distributing confidential rate data broadly.
Evidence: nist-ai-rmf, nist-ai-rmf-core, fmc-service-contracts
Implementation sequence
Prove one forecast-to-buy decision loop before scaling a freight-intelligence platform
Start with one corridor whose rate basis, demand, quotes, contracts, bookings, invoices, and service outcomes can be reconstructed. It needs enough decision frequency to evaluate, material exposure, and an accountable buyer. Avoid choosing a lane only because index history is abundant if the organization never buys on that basis. Include ordinary and disrupted periods in historical tests, then run live shadow mode.
Separate workstreams while connecting evidence. Data teams build vintage and ontology controls. Forecast teams compare transparent baselines and calibrated candidates. Procurement defines feasible actions and risk tiers. Legal and finance review formula and contract records. Operations defines capacity and service outcomes. Security controls commercial data and tools. A custom logistics workflow can connect those roles without giving the model uncontrolled booking or supplier communication.
Progressive authority should stop at decision support unless separately proven and authorized. The first release can publish a forecast range, basis diagnostics, normalized quote comparison, and scenario worksheet. The buyer records the decision. Later releases can draft sourcing events or contract calculations for approval. Automatic commercial commitment is a materially different risk and should never be smuggled into scope as a convenience feature.
- Define the decision and target
Name corridor, equipment, rate scope, horizon, forecast target, buyer action, eligible volume, timing, capacity constraints, success measures, and prohibited autonomous actions.
- Create a canonical rate and event ontology
Normalize geography, equipment, cargo, components, roles, validity, capacity, award, booking, sailing, invoice, and service while preserving original wording.
- Build point-in-time data
Store observation, publication, ingestion, correction, and decision timestamps. Reconstruct index, trade, PortWatch, quote, contract, and internal feature vintages.
- Measure the current buying baseline
Capture tender labor, comparable bids, components, allocation, booking acceptance, rollover, transit, accessorials, invoice variance, service, and exceptions.
- Benchmark simple and advanced forecasts
Use rolling origins and untouched tests. Compare naive, seasonal, statistical, machine-learning, and analyst baselines on point, interval, tail, stability, and cost.
- Design a frozen decision simulation
Specify feasible actions, lead times, transaction and switching costs, volume limits, capacity, service, regret, and a non-hindsight baseline before testing value.
- Backcast any index formula
Apply exact series, lag, average, adjustment, floor, ceiling, rounding, correction, and fallback. Measure basis and operational reproducibility.
- Run shadow procurement
Publish scenarios without changing awards. Let buyers record intended actions and reasons, then compare with rate, quote, capacity, and service outcomes.
- Launch one governed lane
Use buyer approval, materiality tiers, source lineage, data and model versions, capacity checks, security, rollback, and an alternate policy.
- Expand through replicated evidence
Revalidate every corridor, equipment, index, mechanism, and horizon. Monitor calibration, basis, supplier performance, regret, incidents, and total cost.
Evidence: nist-ai-rmf, sse-scfi-faq, baltic-fbx-guide, fmc-service-contracts
Operating policy
Choose the mechanism by uncertainty, basis, capacity, and execution need
Lead with AI forecasting when the decision benefits from scenario timing and the organization has reliable vintage data, a feasible action set, and a buyer who can interpret uncertainty. Keep a transparent baseline and abstain on sparse, transformed, or out-of-regime lanes. Do not use the model output as a purchase-order value. Its job is to shape risk posture, not manufacture a counterparty.
Lead with index-linked buying when a representative licensed benchmark exists, historical and stress basis are acceptable, both parties can reproduce the formula, and capacity and service are contracted separately. Index a defined component rather than an ambiguous total. Use floors, ceilings, lag, and extraordinary-event terms only after showing how they allocate risk. Maintain successor and dispute processes.
Lead with forwarder or NVOCC quotes when the business needs executable coverage, nonstandard geography or cargo, tactical flexibility, routing alternatives, local service, or a counterparty to manage documentation and exceptions. Normalize the full offer and verify role, status where applicable, capacity, validity, exclusions, and service history. Use indices and forecasts as context rather than substitutes for terms.
Use a portfolio when lanes differ. Benchmark-aligned corridors may support an indexed core. Launch-critical or service-sensitive volume may justify committed quote or service-contract coverage. Uncertain demand may remain tactical within a risk limit. Forecasts can adjust review intensity and scenario allocation without mechanically shifting the whole portfolio. The policy should name caps, triggers, approvals, and recovery paths before volatility arrives.
FAQ
Can AI accurately predict ocean freight rates?
AI can produce useful forecasts for a precisely defined index, quote, spread, or regime, but accuracy varies by lane, horizon, sample, regime, features, and evaluation design. No published model proves universal accuracy. Require point-in-time backtests, naive baselines, calibrated intervals, tail tests, and a decision simulation on the organization's actual rate basis before production use.
Is an index-linked ocean contract always cheaper than a fixed contract?
No. Indexation changes the contractual response to a benchmark; it does not guarantee a lower result. Cost depends on the market path, formula, adjustment, lag, floor, ceiling, basis, eligible volume, surcharges, capacity, and service. Evaluate historical and stress scenarios for both parties and select the mechanism for transparency and durable risk sharing rather than retrospective lowest cost.
What is basis risk in ocean freight indexation?
Basis risk is the difference between the referenced benchmark and the shipper's actual commercial exposure. It can arise from ports, inland scope, equipment, cargo, tenor, rate components, timing, carrier or service mix, and methodology. Measure aligned historical residuals and tails, define a review rule, and preserve a successor process for benchmark change or cessation.
Why not use a global container index to forecast every lane?
A global composite summarizes weighted corridors and can provide macro context. Individual directions, ports, equipment, supplier segments, and seasons can diverge. Use the most specific qualified target and test whether the global feature adds held-out value. Do not translate the global level directly into a local procurement rate without a proven, monitored basis relationship.
How should a shipper compare forwarder quotes fairly?
Normalize geography, equipment, cargo, dates, validity, capacity, routing, transit, base rate, every surcharge, terminal and inland scope, documentation, free time, accessorials, taxes, exclusions, provider role, and service. Flag missing terms rather than silently imputing them, let providers confirm material fields, and evaluate total executable cost with capacity and performance.
Which metrics should an ocean-rate forecast report?
Report point error, directional or regime performance, empirical interval coverage and width, tail miss severity, calibration, stability, data freshness, and results by corridor and horizon. Then separately report simulated decision regret, capacity execution, service, and total decision cost. A single accuracy or savings percentage is too weak for procurement governance.
Can the model automatically book freight when it predicts an increase?
That should not be the starting design. A forecast is not an executable rate, and booking requires authority, counterparty, commercial terms, capacity, cargo, routing, timing, and operational checks. Begin with shadow scenarios and buyer approval. Any later automation needs separate authorization, typed constraints, independent price and capacity validation, audit, rollback, and exception recovery.
How often should an index-linked ocean formula reset?
There is no universal cadence. Weekly, monthly, or quarterly resets create different responsiveness, noise, lag, workload, and risk sharing. Backcast the exact observation window and effective lag through stable and disrupted periods, test calculation and invoice timing, and choose a cadence both parties can reproduce that matches the underlying exposure.
A fictional consumer-goods sourcing investigation
A retailer discovers that its forecast was right while its procurement conclusion was wrong
Consider a fictional retailer importing containerized home products from several Asian origins into North America. Its data team builds a route-index model that correctly signals a rising period more often than the existing moving-average baseline in an initial retrospective test. Leadership concludes that the model should trigger earlier fixed awards. A buyer objects: the target is a Shanghai base-port spot index, while much of the cargo originates inland, uses several load ports, includes destination and security components, and depends on named-account capacity. Statistical success and commercial exposure are not the same object.
The team rebuilds the test with point-in-time data and a rate ontology. It removes final invoices, retrospective event tags, and revised features unavailable at forecast cutoff. It separates base ocean, indexed surcharges, fixed charges, inland, terminal, and accessorials. It also adds quote validity, allocation, booking acceptance, rollover, routing, transit, and invoice outcomes. The forecast still provides useful scenario information, but its apparent error advantage narrows and varies by corridor.
Procurement then tests three mechanisms. A qualified index formula is backcast on the most aligned high-volume corridor, including lag, collar, fixed adjustment, and basis. Forwarder quotes remain primary for secondary ports, tactical volume, and nonstandard scope. The model supplies calibrated ranges and flags lanes where rising risk and launch-critical demand justify buyer review. It never generates a bookable rate or accepts a quote.
The production result is a mixed decision stack rather than an AI buying bot. The indexed corridor receives a mechanism both parties can reproduce, quote lanes receive normalized executable comparisons, and forecasts shape scenarios and review priority. The team measures calibration, basis, quote comparability, capacity, service, total freight, decision regret, and procurement effort separately. This is an illustrative composite, not a client engagement, savings claim, or current market forecast.
- The original model forecasted an index rather than the retailer's executable all-in exposure.
- Point-in-time reconstruction removed revised and outcome data from the historical test.
- Indexation was limited to a corridor whose scope and basis could be governed.
- Forwarder quotes remained essential for executable capacity and complex lanes.
- The model influenced scenarios and review priority while buyers retained commercial authority.
Research method, market-data boundaries, and scenario disclosure
This investigation was researched against primary and authoritative materials available on August 30, 2026: UN Trade and Development's Review of Maritime Transport 2025 and connectivity resources; the WTO March 2026 outlook; Shanghai Shipping Exchange SCFI documentation; Baltic Exchange FBX methodology; IMF PortWatch research; Federal Maritime Commission service-contract and OTI materials; and NIST's AI Risk Management Framework. It also uses peer-reviewed freight-rate forecasting research to show that performance depends on model, route, sample, target, and evaluation. No current spot rate, forwarder quote, carrier offer, savings percentage, or procurement return is invented or implied. Regulatory examples are U.S.-specific illustrations, not legal advice.
Research ledger
Sources and further reading
- Review of Maritime TransportUN Trade and Development
Official flagship-series page stating maritime transport's share of international goods trade by volume.
- Review of Maritime Transport 2025: Staying the Course in Turbulent WatersUN Trade and Development · 2025-09-24
Primary source for 2024 seaborne volume and ton-mile growth, rerouting, disruption, rates, fleet, and ports.
- Global Trade Outlook and Statistics — March 2026World Trade Organization
Official WTO outlook reporting 2025 merchandise trade growth and the 2026 baseline and scenarios.
- Shanghai Containerized Freight Index: Compilation and Publication FAQShanghai Shipping Exchange Institute (in Preparation)
Official SCFI scope, units, included and excluded charges, panel data, calculation, base, and cadence.
- Freightos Baltic Global Container Index GuideBaltic Exchange Information Services Limited
Official benchmark guide documenting the FBX family, input data, determination, governance, and policies.
- Baltic Exchange Indices — ContainersBaltic Exchange
Official overview of weekly 40-foot-container spot-rate indices across twelve trade lanes.
- Indexing OverviewXeneta
Provider documentation used only to describe commercial index-contract mechanics, not to prove savings.
- IMF and University of Oxford Launch PortWatch PlatformInternational Monetary Fund · 2023-11-15
Official launch description of satellite-based maritime disruption monitoring and simulation.
- Nowcasting Global Trade from SpaceInternational Monetary Fund
IMF working paper describing PortWatch indicators, revisions, and conceptual limitations.
- Liner Shipping Connectivity Index Data InsightsUN Trade and Development Data Hub · 2026-06-26
Official current description of LSCI meaning, six components, revised reference, and data.
- How to File Service ContractsU.S. Federal Maritime Commission
Official U.S. overview defining ocean service contracts, commitments, service, filing, and records.
- Ocean Transportation IntermediariesU.S. Federal Maritime Commission
Official explanation of ocean freight forwarder and NVOCC roles, tariffs, contracts, and arrangements.
- U.S. Court of Appeals Issues Decision in Case on Demurrage and Detention Billing PracticesU.S. Federal Maritime Commission · 2025-11-20
Current FMC update illustrating that accessorial governance extends beyond base ocean rates.
- Artificial Intelligence Risk Management FrameworkNational Institute of Standards and Technology · 2023-01-26
Official entry point for NIST's voluntary AI RMF and trustworthiness resources.
- AI RMF CoreNational Institute of Standards and Technology AI Resource Center
Official presentation of Govern, Map, Measure, and Manage functions.
- Forecasting Container Freight Rates for Major Trade Routes: A Comparison of Artificial Neural Networks and Conventional ModelsMaritime Economics & Logistics · 2021-06-01
Peer-reviewed comparison of ARIMA, VAR or VEC, and neural-network approaches using CCFI routes.
- A Comparative Evaluation of Machine Learning Approaches for Container Freight Rates PredictionThe Asian Journal of Shipping and Logistics · 2025-06-01
Open peer-reviewed comparison of decision tree, random forest, Prophet, and LSTM approaches.
- Machine Learning in Freight Rate Forecasting: A Systematic Literature ReviewMaritime Economics & Logistics · 2026-03-02
Peer-reviewed systematic review of machine-learning freight-rate forecasting.
- Parsimonious Decomposition-Based Model Search Engine for Monthly Shipping Freight Rate Interval ForecastingApplied Soft Computing · 2026-07-01
Peer-reviewed shipping-rate research focused on interval rather than point forecasts.
Build an ocean buying system that preserves uncertainty and execution truth.
bizz designs freight data, AI forecast, index calculation, quote normalization, sourcing, and exception workflows around auditable commercial decisions.
Explore logistics software