Evaluating LLM Capabilities for Commodity Classification

Executive Summary

  • Commodity classification is the act of determining what category an item falls within for the purpose of export control.

  • Commodity classification is a common task at the Bureau of Industry and Security (BIS), which administers the U.S. government’s dual-use export controls.

  • Automating commodity classification could free up BIS staff to process export licenses more rapidly, and give BIS experience in building in-house AI tools.

  • The Institute for AI Policy and Strategy (IAPS) tested large language models (LLMs) for this task and the preliminary results were promising, but additional testing with more realistic inputs and a wider variety of items is needed.

  • BIS should establish a pilot program to test LLMs for commodity classification.

What is Commodity Classification?

Commodity classification is the task of reviewing information about an item (a physical item, a piece of software, or an intangible technology), and determining what Export Control Classification Number (ECCN) it falls under on the Commerce Control List. ECCNs always have at least five characters: a number, indicating which CCL category the item falls under; a letter, indicating which product group the item belongs to; and three more numbers uniquely identifying that type of item. ECCNs also sometimes have subparagraphs specifying different sub-categories of one ECCN.

Benefits of LLMs for Commodity Classification

Automating commodity classification could allow BIS to review licenses faster while continuing to protect national security.

  • In April 2026, a survey of exporters by the Center for Strategic and International Studies reported that export license applicants were experiencing increased delays and reduced transparency from BIS in adjudicating license applications.

  • In his testimony to the House Foreign Affairs Committee on July 14, 2026, Under Secretary of Commerce for Industry and Security Jeffrey Kessler stated that BIS was applying increased scrutiny to licenses.¹

  • Automating some of the work licensing officers carry out, like commodity classification, could allow BIS to process licenses more rapidly without sacrificing rigor. Processing licenses more rapidly would allow more exports to low-risk destinations and advance Pillar III of the AI Action Plan while still protecting national security.

Automating commodity classification could enable BIS to build in-house AI expertise to take on more ambitious projects in the future.

  • LLMs are rapidly improving in capabilities, and the administration has prioritized rapidly deploying AI for government

  • Successful AI deployment in government will depend not only on having access to advanced models, but also on applying domain expertise to designing processes and scaffolding around the models and identifying the most promising use cases.

  • Commodity classification is a low-cost, closed-ended application that could give BIS practice deploying AI models internally and build organizational buy-in for more ambitious projects to accelerate licensing, rulemaking and enforcement.

Evaluating LLM Performance at Commodity Classification

The acceptable level of errors in commodity classification is zero, so LLMs used for the task must either not make errors or the outputs must be reviewed by a human. Because verifying a commodity classification requires the same amount of effort as performing it, human review of AI outputs does not save labor compared to humans directly performing the task. Thus, for LLMs to be useful at this task, they must be at least as reliable as a human licensing officer.

IAPS used qualitative methods to evaluate LLM performance. Specifically, IAPS provided some example questions to the Claude Opus 4.8 model to validate whether models are capable of commodity classification at all and attempt to discern qualitative patterns. The questions and answers received are in the Appendix to this brief. The model was provided with no scaffolding other than an “ecfr” tool that it could use to access the Code of Federal Regulations API.²

Patterns in model behavior when classifying commodities

The model sometimes leaned on hints provided in the question. In one response, the model noted that the parameters provided in the item description matched the thresholds asked for in a specific ECCN. This reinforces the need to conduct benchmarking with real BIS data or more realistic item descriptions including large amounts of extraneous information.

The model successfully searched across the Export Administration Regulations and combined different parts of the rules. In one response, the model combined information from the Commerce Control List and Part 772 of the Export Administration Regulations (containing definitions) to successfully classify an item.

The model successfully classified using multimodal information and context beyond the description. In one response, the model correctly identified the item in an image as an NVIDIA Vera Rubin AI accelerator board when provided only the image. In another response, the model correctly identified that an ASML product was a controlled extreme ultraviolet (EUV) lithography machine based on the name and model name of the product.³

The model successfully applied general scientific knowledge and mathematics. In one response, the model used a physical description of part of the components of an item to determine that it was a specific type of item. The model also carried out unit conversions accurately.

Directions for future research

In order to rigorously evaluate model performance, asking one-off questions of a commercial model with minimal scaffolding is insufficient. For a future benchmark constructed by BIS or an outside organization, the following considerations should apply:

The input provided to the model should resemble real inputs as much as possible. Rather than providing a short item description, the model should be provided manufacturer datasheets, letters describing the end-use, and other information in formats and of types similar to that typically provided to BIS as part of a license application or commodity classification request.

The dataset should prioritize covering a diverse range of items and as many unique edge cases as possible. For items requiring licensing applications that BIS regularly processes, there likely is no need for commodity classification assistance because the items have already been reviewed many times and their classification has been clearly established. Items that require commodity classification will be more ambiguous or rarer than typical ones.

The benchmark should test models from multiple providers at multiple levels of capability. LLMs vary in their capabilities across providers and model generations. Newer models are generally more capable, but also more expensive to run. Older models may have sufficient performance for this task without excessive token costs.

Recommendations

BIS should create a pilot program to test LLMs for commodity classification in a more realistic environment.

  1. BIS should create a pilot program using the pilot program exemption in Section 4(a)(i) of OMB Memorandum M-25-21, “Accelerating Federal Use of AI through Innovation, Governance, and Public Trust.” BIS should seek certification for this program from the Department of Commerce Chief Artificial Intelligence Officer.⁴

  2. For the pilot program, BIS should use AI tools that it already has access to and that meet the requirements for handling export license information, which is a form of Controlled Unclassified Information (CUI). ChatGPT Enterprise is available at FedRAMP Moderate and would be suitable for this purpose, although BIS should also explore alternative vendors.

  3. BIS should also coordinate with the Center for AI Standards and Innovation, also housed within the Department of Commerce, on best practices for deploying AI securely in government.

  4. BIS should use a pilot program to explore the following questions:

    a. How accurate are the AI models available to BIS at commodity classification tasks?

    b. Which types of errors are most common? Can external scaffolding or data cleaning address these error types?

    c. What other use cases at BIS could benefit from sufficiently reliable AI deployments?

Appendix: Tested Questions and Responses

Endnotes

  1. Testimony of Under Secretary Jeffrey Kessler to the House Foreign Affairs Committee, July 14, 2026 https://www.youtube.com/watch?v=XcKNV5ZqyAg&t=2890s (57:42)

  2.  eCFR.gov blocks access by bots.

  3. The model did not appear to use web search to do so, suggesting that information about that item was contained in the model weights.

  4. Export control commodity classification would ordinarily be a high-impact use case, per 6(h), triggering additional requirements described in 4(b). However, BIS can leverage the provisions of section (4)(a)(i) to exempt a pilot program from such requirements.

  5. Appendix (Item description 3): This was an unintentional typo by the researcher (the actual machine is an ASML TWINSCANEXE:5000), but the model nevertheless classified the item correctly.

Next
Next

The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems