Head to head · updated June 22, 2026
We asked ChatGPT, Claude, and Gemini to classify five real invoice lines.
Then we checked their answers against broker-verified HTS codes. The chatbots are smarter than you'd expect — and still not something you can file. The same contradictory line got opposite codes from two of them, and not one of the fifteen answers was anchored to the CBP ruling that governs it. Here is the side-by-side, with the methodology in full.
answers cited the governing CBP CROSS ruling
returned opposite codes on the same line — with no warning
is where the chatbots and a broker fully agree: the easy one
How we ran this
The test, in full — so you can repeat it.
The prompt (identical every time)
“I'm importing the following into the United States: [line item]. What is the correct 10-digit HTS (Harmonized Tariff Schedule) classification code?”
What we tested, June 22, 2026
ChatGPT (chatgpt.com, default model), Claude Opus 4.8 (claude.ai, web search on), and Gemini (gemini.google.com, Pro). Each line was a fresh chat. Answers were captured verbatim.
The answer key
Ground truth is CrossCode's broker-verified classifications — each with the CBP CROSS ruling, GRI rule, and the duty rate behind it.
What we scored
Did the code match? Did it cite the governing CBP ruling? Did it flag a genuine ambiguity instead of guessing? This is a point-in-time test — these tools change weekly, and your results may differ.
The line that says it all
One contradictory line. Two tools. Opposite codes. Zero warnings.
We gave each tool: Men's cotton T-shirt, woven, short sleeve. In the tariff that phrase contradicts itself — a “T-shirt” is knitted (heading 6109); “woven” forces heading 6205. A safe classifier flags the conflict. Here's what came back.
Picked knitted. Ignored “woven” entirely. No warning.
Picked woven. Opposite code to ChatGPT. Also no warning.
Flagged the conflict. Explained both, asked for the fabric spec.
Flagged + provisional + ruling. Conflict flagged at 52% confidence, 6109 alternative + CROSS NY N121394, sent to a broker.
Either chatbot, taken at its word, files a different code on the same goods — and one of them is wrong on the entry. Neither told you there was a decision to make. That is the whole risk in one line.
All five lines
The full results.
On the easy line, everyone's right. The gap opens exactly where the money and the audit risk live: ambiguity, legal notes, and the 10-digit leaf.
Men's cotton T-shirt, woven, short sleeve
Self-contradictory line. In the tariff, a “T-shirt” is knitted (heading 6109); “woven” forces a different chapter (6205). The two facts can't both be true — a classifier should flag it, not silently pick one.
Ignored the word “woven” and returned a knitted code. No flag that the description contradicts itself. The only “source” shown was a sponsored ad.
Caught the contradiction immediately, laid out knitted 6109 vs woven 6205 with duty rates, and refused to pick without the fabric spec. Cited trade blogs, not a CBP ruling.
Did the opposite of ChatGPT — honored “woven,” committed to a woven-shirt code, and never warned that “woven T-shirt” is a contradiction. No CBP ruling.
Flagged the woven-vs-knitted conflict (confidence 52%), returned a provisional code plus the 6109 alternative, and routed it to a licensed broker before filing.
Stainless steel kitchen knife set
A “set” triggers special rules: the 8211.10 sets provision, GRI 3(b) essential character, and Chapter 82 Note 3 — which can pull the whole set into a different heading and duty rate.
Reached the sets provision and flagged the heading-8215 cutlery rule — but cited tariff-lookup blogs, not a CBP ruling, and missed the stainless-handle leaf.
Surfaced the Chapter 82 Note 3 set-duty trap, distinguished assorted vs uniform sets, and asked for contents and origin. No CBP ruling anchored.
Landed on the sets provision and explained the blended set-duty rate, but cited no specific CBP ruling and didn't reach the stainless-handle leaf.
Resolved via GRI 3(b), surfaced the Chapter 82 Note 3 caveat, committed a code with the governing ruling, and had a broker confirm set composition.
USB-C charging cable, braided sheath
Heading 8544.42 is easy; the 10-digit leaf is not. It splits on telecommunications-use (.2000) vs other (.9000) — the kind of statistical-suffix call CBP rulings actually turn on.
Right heading, reasoned about telecom vs power — but cited a trade-finance blog and gave no governing CBP ruling.
Committed to a leaf, and its search even surfaced a real CBP ruling (NY N007536) — but the answer wasn't anchored to it.
Leaned to the telecom leaf as “most common,” the opposite of the charging-use call. Conditional, with no specific ruling cited.
Resolved the telecom-vs-charging split toward the charging-use leaf, with the supporting ruling and a confidence score, broker-reviewed.
Glazed ceramic dinner plates
“Glazed” tells you nothing — both relevant headings are glazed. The code turns on the ceramic body: porcelain/china (6911) vs stoneware/earthenware (6912), plus size and value breakouts.
Correctly chose non-porcelain (6912 over 6911) and hedged on the leaf — but cited a tariff-lookup site, not a CBP ruling.
Refused to invent a 10-digit (“a guess dressed up as an answer”) and mapped the 6911-vs-6912 fork. No CBP ruling.
Same 6911-vs-6912 fork; asked for body type, dimensions and value before committing. No CBP ruling.
Classified as non-porcelain ceramic with the governing ruling and the 6911 porcelain alternative surfaced for the broker to confirm.
Bamboo cutting boards, food-prep use
The clean one. An unambiguous, well-documented good with an eo nomine subheading (4419.11, added in the 2022 restructuring). Everyone should get this right.
Correct, including the post-2022 subheading — but the cited source was a tariff-lookup site, not a CBP ruling.
Correct, with clean GRI-1 reasoning and a sharp Section 301 duty analysis. Cited trade blogs, not a CBP ruling.
Correct, and noted CBP is the ultimate authority — but cited no governing ruling.
Correct code with the supporting ruling, a confidence score, and a broker sign-off — the same audit packet every line gets.
What every chatbot missed
Four gaps that no prompt fixes.
Not one answer cited the ruling CBP would expect.
Across all fifteen runs, the citations were trade blogs (Flexport, Drip Capital, Gaiadynamics) and tariff-lookup tools (TariffLens, HTS Hub, tariffnumber.com), or a generic wave at “CBP rulings.” One model's web search surfaced an actual CROSS ruling — and then didn't build the answer on it. A CBP entry needs the governing ruling, not a blog that paraphrases one.
They won't commit — and they make you the expert.
On every genuinely ambiguous line, the stronger models correctly asked us for fabric construction, set contents, or ceramic body type. That's the right instinct. But it hands the hard part back to you: you already have to know what to ask, and what you get back is a paragraph, not a filing.
Nothing is broker-reviewed, so nothing is filing-defensible.
You cannot put “ChatGPT said so” on an entry. Reasonable care under 19 U.S.C. § 1484 stays on the filer. No chatbot output carries a licensed broker's name, a confidence score tied to a review threshold, or an audit packet you can attach to the entry file.
It's not repeatable, and it depends on which tier you're on.
The logged-out, free ChatGPT returned the most dangerous answer of the test and hit a login wall mid-session. Paid tiers behave differently; the same prompt yields different depth run to run. A compliance process needs the same input to produce the same defensible output every time.
What CrossCode returns instead
The same line — as something you can actually file.
A committed 10-digit code
Not a paragraph of options — the actual statistical suffix you file.
The governing CBP CROSS ruling
The specific ruling number CBP itself has issued for a materially similar good.
The GRI rule applied
So an examiner can follow the reasoning instead of taking the code on faith.
A confidence score
And an automatic defer-to-broker when a line is ambiguous — the woven T-shirt never ships silently.
A licensed broker's sign-off
Every code reviewed and accepted by a named CHB before it leaves your shop.
The whole invoice in under 60 seconds
Not one line at a time in a chat window — every line, with an exportable audit packet.
To be fair to the chatbots.
This isn't a story about dumb AI. Claude and Gemini reasoned about GRI rules, Chapter legal notes, and Section 301 duty stacking at a level that would hold its own against a lot of humans. On a clean, well-documented good they nail the code. The point is narrower and more important: a chatbot is built to give you a good answer in a conversation. Customs classification needs a defensibleone — a committed code, tied to the CBP ruling behind it, with a licensed broker's name on it, produced the same way every time. That is a different job, and it's the one CrossCode was built for.
See it on your own invoices.
30-minute call. We classify a real invoice from your shop while you watch.