AI · Business Data · Distributors

Without These Business Tables, AI Can Only Talk in Platitudes

For AI to understand a real business, it must see verifiable sales, inventory, customer, and account data. Without a data foundation, its answers never rise above common sense.

Many business owners run into the same problem the first time they ask AI to analyze their business.

Ask “how do I optimize inventory,” and it says to improve turnover and reduce overstock. Ask “how do I manage customers,” and it says to tier them, maintain relationships, and watch repeat purchases. Ask “how do I raise profits,” and it says to control costs and optimize the product mix.

It all sounds right, yet it amounts to saying almost nothing.

At this point many people conclude: AI doesn’t understand the FMCG business.

In fact, the more common reason is this: AI has never actually seen your business.

It doesn’t know how many products you carry, doesn’t know how a pack, a case, and a multipack convert into one another, doesn’t know that the same store has three different names in your system, and doesn’t know whether yesterday’s negative entry was a return or a reversal. If all you give it is a summary sheet — or just a one-line question — of course all it can offer is common sense.

For a distributor to move AI from “good at chatting” to “good at working,” you don’t need to build a data platform first, and you don’t need to replace your ERP. The first step is to make a few of the most basic business tables clearly explained, mutually consistent, and traceable.

This article walks through five tables: products, customers, sales, inventory, and accounts receivable.

They are not an IT department project; they are the five ledgers through which an owner sees the business clearly.

1. More Data Isn’t Better — First Make Sure It Reconciles

Many companies are not short on data.

There are orders in the ERP, stock in the warehouse, receivables in finance, visit records on the sales reps’ phones, and policy documents in the manufacturer’s group chat. The problem is that these materials are scattered across different places, with inconsistent names, units, dates, and definitions.

The four most typical situations:

  1. Same item, different names: the same beverage is called “500ml×24” in the purchasing table, “large bottle” in the sales table, and yet another internal shorthand in the warehouse table;
  2. Same store, multiple identities: one store has its business-license name and its storefront name, and the sales rep has given it a nickname on top of that;
  3. Unit chaos: some tables count in cases, some in packs, and the return records count in bottles;
  4. Time misalignment: sales are booked by order date, finance by payment date, inventory by the midnight snapshot of the day — and the owner subtracts the three tables from one another directly.

Until these problems are solved, the faster AI calculates, the faster it gets things wrong.

So the first standard for data preparation is not “intelligence” but four sentences:

  • Every object has a unique code;
  • Every field has a clear meaning;
  • Every number has a unit and a time;
  • Every result can be traced back to the original document.

2. Table One: The Product Table — First Settle “What Exactly Are You Selling”

The product table is the starting point of all analysis.

At a minimum, it should include:

Field Why it’s needed
Product code Prevents confusion from differing names, abbreviations, and spec notations
Product name For human readability
Brand, category For viewing the product mix
Specification States how many packs per case and how many bottles per pack
Base unit, sales unit Prevents mixing up cases, packs, and bottles
Tax-inclusive purchase price, reference selling price For basic profit analysis
Shelf life For near-expiry risk
Launched/discontinued status Prevents treating discontinued products as growth opportunities

The most common trap in the product table is using the “name” as the code. Names can change; codes must not be changed casually. Otherwise, each time a product’s name changes, the historical data looks as if a new product has appeared.

The second trap is keeping only the purchase price and selling price without stating dates. After the manufacturer adjusts prices, today’s purchase price cannot simply be applied to sales from six months ago.

If you don’t yet have a complete cost history, state it plainly first: “This analysis uses the reference purchase price as of July 26, for estimation only; it does not represent actual historical costs.” That is more reliable than pretending to be precise.

3. Table Two: The Customer/Store Table — First Confirm “Who Did You Sell To”

The customer table is not just a contact list.

The fields genuinely useful for a distributor’s business analysis include:

  • Customer code;
  • Store name and region;
  • Channel type;
  • Customer tier;
  • Person in charge or sales rep;
  • Agreed payment terms;
  • Whether cash on delivery;
  • Date of the most recent valid transaction;
  • Status: operating, closed, or pending verification.

Why do you need a customer code?

Because “Lao Wang Supermarket,” “Boss Wang’s Store,” and “South Side Lao Wang Convenience” may all be the same store. If AI treats them as three stores, repeat purchases, receivables, and sales volume will all be calculated wrong.

Why do you need a status field?

Because a store that hasn’t ordered for 60 straight days may be a lost customer — or may have already closed down. Data can only tell you “no orders”; it cannot judge the reason on the sales rep’s behalf.

Customers’ phone numbers, national ID information, personal WeChat accounts, and the like should not all be uploaded just for analytical convenience. Analyzing store sales usually requires only the customer code, type, region, and transaction data. If the job can be done without personal information, don’t carry an extra unit of risk.

4. Table Three: Sales Detail — Don’t Give AI Just a Monthly Total

If all you provide is a summary sheet saying “this month’s sales: 3 million yuan,” AI will find it very hard to spot problems.

Useful sales detail must at least come down to:

  • Date;
  • Order number;
  • Customer code;
  • Product code;
  • Quantity and unit;
  • Sales amount;
  • Return quantity and amount;
  • Free-gift or promotion flag;
  • Sales rep;
  • Document status.

Why keep the order number?

Because analysis results must be traceable. If AI says a certain customer’s returns this month are abnormal, finance must be able to follow the order number back to the original document, rather than spending another half day searching by hand.

Why must returns be listed separately?

If the system records returns as negative numbers and you don’t explain that, AI may delete them as data errors, or feed the negatives straight into the average selling price, distorting the conclusions.

Why include document status?

Drafts, voided documents, approved documents, and reversed documents must not be mixed together. What the owner needs to see is the business that actually happened, not every record that ever appeared in the system.

5. Table Four: The Inventory Table — One Snapshot Cannot Tell the Whole Story

The inventory table needs at least:

  • Snapshot date and time;
  • Warehouse;
  • Product code;
  • Available stock;
  • Locked stock;
  • Quantity in transit;
  • Batch and expiry dates;
  • Most recent inbound date;
  • Most recent outbound date;
  • Inventory cost basis.

Looking at stock quantity alone makes misjudgment easy.

Take the same 500 cases of goods: one product sells 100 cases a day and clears in five days; another sells 5 cases a day and will sit for a hundred. High inventory is not necessarily dangerous — the key is to view it together with sales velocity, expiry dates, and the purchasing cycle.

“System stock” is also not the same as “sellable on the floor.” Damaged goods, pending returns, locked stock, and stocktaking discrepancies can all be occupying the number. AI can spot anomalies, but the warehouse keeper has to confirm what’s physically there.

6. Table Five: The Receivables Table — Beyond Amounts, Look at Time and Promises

The receivables table should at minimum include:

  • Customer code;
  • Original document number;
  • Amount receivable;
  • Amount received;
  • Amount outstanding;
  • Invoice or shipment date;
  • Agreed due date;
  • Most recent payment date;
  • Person in charge;
  • Dispute or reconciliation status.

Many owners look only at “who owes the most,” but the largest amount is not necessarily the most dangerous.

A long-term chain customer whose balance is not yet due under the contract, and a smaller-amount customer who has repeatedly promised to pay and then broken the promise — these are not the same kind of risk.

What AI can do is lay out the days overdue, the amounts, the changes in payment behavior, and the gaps in the records. Whether to extend credit, whether to suspend supply, and how to communicate must still be decided by the person in charge, weighing the customer relationship, the contract, and market conditions.

7. A Complete Demonstration: Why the Same Sales Data Yields Such Different Results

First, look at an incomplete dataset:

Date Customer Product Quantity Amount
July 20 Lao Wang’s Store A beverage 10 460
July 21 Boss Wang Supermarket A beverage 500 -2 -92

Without further explanation, AI has no way to know:

  • whether the two customers are the same store;
  • whether the quantity unit is cases or packs;
  • whether the second row is a return, a reversal, or an error;
  • whether the products are the same specification;
  • whether the amounts include tax;
  • whether the documents have been approved.

After master data and field descriptions are added:

Order No. Customer Code Product Code Quantity Unit Sales Amount Return Amount Status
S072001 C018 SKU036 10 Case 460 0 Approved
R072101 C018 SKU036 0 Case 0 92 Approved

Then tell AI:

  • C018’s storefront name is “Lao Wang Convenience Store”;
  • SKU036 is 500ml×24 bottles;
  • return documents are counted separately under return amount;
  • this run is to calculate net sales amount and net sales volume;
  • do not analyze return causes — only list the items awaiting the sales rep’s confirmation.

Only then can it conclude:

  • net sales volume: 8 cases;
  • net sales amount: 368 yuan;
  • 2 cases returned by the same customer;
  • insufficient information on the cause of the return; the sales rep needs to supplement it.

This is not AI suddenly getting smarter. It is you finally giving it materials and definitions it can work with.

8. Sort Data into Red, Yellow, and Green Before Deciding What Can Go into AI

A distributor doesn’t need to build a sophisticated data security system from day one, but there must at least be a simple tiering.

Green: usable for everyday analysis

  • De-identified product codes;
  • Store codes and channel types;
  • The necessary fields from sales, inventory, and receivables;
  • Business records containing no personal contact information;
  • Already-public product materials and company policies.

Yellow: use only after processing

  • Customer contacts’ names and phone numbers;
  • Personal evaluations of employees;
  • Specific pricing policies and rebates;
  • Contract excerpts;
  • Sales reps’ voice recordings and store photos.

For this class of material, first judge whether it is truly needed, then de-identify, excerpt, and obtain authorization.

Red: never enters public AI by default

  • National ID numbers, bank card numbers, passwords, and verification codes;
  • Complete, unpublished customer lists;
  • Trade secrets and unauthorized contracts;
  • Sensitive information that can directly identify individuals;
  • System accounts, keys, and access tokens.

The Chinese national standard GB/T 43697—2024 provides general rules for data classification and grading. A distributor need not copy an entire professional regime wholesale, but the three principles — classify first, then use, and take only the minimum needed — are non-negotiable.

9. Get the Data to “Good Enough” in Seven Days — Don’t Chase Perfection in One Go

Day one: pick one scenario, such as the daily report, inventory anomalies, or receivables follow-up.

Day two: list the fields this scenario truly needs; don’t export the entire ERP database.

Day three: unify product codes, customer codes, dates, and units.

Day four: spell out the special definitions — returns, voided documents, free gifts, in-transit, locked stock, and so on.

Day five: spot-check ten records against the original documents.

Day six: apply the red/yellow/green tiering; delete unnecessary personal information and sensitive fields.

Day seven: have AI complete one analysis, with the responsible business lead accepting it item by item and logging the errors.

A week later, what you have is not a “data governance project” but a minimal data package that can get concrete work done.

10. How to Use the Distributor AI Data Preparation Checklist

The appendix breaks the five tables down into fields, definitions, common problems, sensitivity levels, and the person responsible for acceptance.

When using it, don’t tick off every item in one sitting. First pick the single scenario this article is meant to solve, then determine:

  • which tables are needed;
  • which fields in each table are essential;
  • which fields are missing;
  • who can confirm them;
  • what information must be de-identified;
  • how the results will ultimately be reconciled against the source system.

If a piece of data has no owner, don’t hand it to AI yet. Because when something goes wrong, no one will be able to explain what actually happened.

Today’s Action

Export the table you use most often from your ERP.

Don’t rush to upload it. First add four lines of explanation: what table this is, the time period it covers, what each row represents, and what units the amounts and quantities are in.

Then randomly sample ten rows and check whether the product codes, customer codes, returns, and document statuses reconcile.

Only when this step is done has AI truly seen your business for the first time.


References

  1. State Administration for Market Regulation and Standardization Administration of China: “GB/T 43697—2024 Data Security Technology — Rules for Data Classification and Grading”.
  2. National People’s Congress of China: “Data Security Law of the People’s Republic of China”, 2021-06-10.
  3. Cyberspace Administration of China and six other departments: “Interim Measures for the Administration of Generative Artificial Intelligence Services”, 2023-07-13.
  4. Internal project research working paper: “Research Notes on AI Adoption by FMCG Distributors”, 2026-07-24.

Next in the series: “I Can’t Write Code — So Why Could I Direct AI to Build ABI?” — being able to do the work yourself and being able to break a big undertaking down clearly are two different things.