Your Company Wants AI. Your Data Is Still Full of Duplicates and Missing Information. Now What?

by | Sep 2, 2026 | Uncategorized | 0 comments

AI Sounds Exciting Until It Starts Reading the Data You Already Have

Artificial intelligence has quickly moved from an experimental technology to something many businesses are actively considering for customer service, sales, finance, operations, marketing and management reporting. A company may want an AI assistant that answers customer questions, predicts demand, identifies sales opportunities, analyses expenses or helps managers make faster decisions. The temptation is to start with the technology because that is the exciting part. Yet there is a less exciting question that often determines whether the project succeeds: what does your existing data actually look like? If the customer database contains duplicate companies, supplier names are entered differently across systems, product codes are inconsistent, important fields are blank and employees maintain their own unofficial spreadsheets, AI does not automatically make those problems disappear. It can instead process unreliable information much faster. Before asking what AI can do for the business, management therefore needs to understand whether the information feeding those tools is accurate, complete, consistent and sufficiently organised to support the decisions the company expects AI to help make.

The Same Customer Appears Five Times, but AI Does Not Know They Are the Same Customer

Consider a company that has dealt with ABC Engineering Pte. Ltd. for ten years. The accounting system records the customer as “ABC Engineering Pte Ltd”, the CRM contains “ABC Engrg”, an old spreadsheet lists “ABC Engineering”, another employee created “ABC Eng Pte Ltd”, and an imported database contains “A.B.C. Engineering Pte. Ltd.” To a human employee who knows the customer, these records obviously refer to the same business. A system may not automatically reach the same conclusion unless matching and data-cleaning processes have been implemented. If management asks an AI tool to analyse customer purchasing behaviour, it could potentially treat these records separately, producing an incomplete picture of the relationship. The company may then underestimate total revenue from the customer, misunderstand purchasing patterns or fail to recognise customer concentration. The AI model may be sophisticated, but the underlying problem is surprisingly ordinary: nobody maintained one reliable customer record.

Duplicate Data Can Turn One Business Relationship Into Several Different Stories

Duplicates become more damaging when information is spread across departments. Sales may use one customer record, finance another and customer service a third. The sales team sees S$500,000 of annual revenue, finance sees S$650,000 and management receives a report showing S$580,000 because each system uses different identifiers and timing. Employees may have learned to work around these inconsistencies because they know which spreadsheet or record to trust. An AI system introduced across the organisation does not automatically inherit that institutional knowledge. If management wants AI to answer questions such as which customers are most profitable, which customers are likely to stop buying or where cross-selling opportunities exist, the organisation first needs confidence that transactions belonging to the same customer are actually connected. Otherwise, the company can produce increasingly sophisticated analysis of fundamentally fragmented information.

Missing Information Creates a Different Kind of Problem

Duplicates are visible once someone looks for them. Missing data can be more subtle. Imagine a customer database containing 20,000 records, but 35% have no industry classification, 20% have no accurate location, some have no assigned account manager and many old records contain outdated contact details. Management wants AI to identify which industries generate the strongest margins and which customers should receive particular marketing campaigns. The technology can analyse the information provided, but it cannot reliably infer every fact the company failed to capture. If important fields are systematically missing, the resulting analysis may represent only part of the customer base. The danger is that a polished dashboard or AI-generated recommendation can look authoritative even when it is based on incomplete information. Businesses therefore need to distinguish between having a lot of data and having enough reliable data to answer a particular question.

AI Does Not Turn Bad Data Into Good Data Simply Because the Output Looks Intelligent

One reason poor data becomes particularly dangerous with AI is presentation. A messy spreadsheet obviously looks messy. An AI-generated summary can sound confident, structured and persuasive. Management may therefore give the output more credibility than the underlying information deserves. Suppose the system concludes that customers in a particular industry generate 25% higher margins than average. The statement sounds precise. But if half the customers have no industry classification and cost information is inconsistent, the apparent precision may be misleading. AI can make analysis easier to consume, but management still needs to understand the quality of the inputs and the limitations of the output. Technology changes how quickly information can be analysed. It does not remove the need to ask whether the information represents reality.

Your Employees May Already Know Which Data Cannot Be Trusted

Before launching an expensive data-cleaning project, management should talk to the employees who use the information every day. They often already know where the problems are. A salesperson may say, “Ignore that customer status field because nobody updates it.” Finance may explain that the product margin report is unreliable because freight costs are allocated manually. Operations may maintain a separate spreadsheet because the official system does not capture certain information correctly. Customer service may know that telephone numbers in the CRM are frequently outdated. These workarounds are valuable clues. They show where employees have stopped trusting the formal system and created their own alternative source of truth. An AI project that ignores these realities may simply automate information that employees already know is unreliable.

The Spreadsheet Nobody Talks About May Be Running Part of the Company

Many businesses have official systems for accounting, customer management, inventory and operations while simultaneously relying on spreadsheets that management barely knows exist. One employee exports data from the accounting system every Monday, cleans it manually and sends a spreadsheet to another department. Another employee keeps a private customer list because the CRM contains duplicates. A manager maintains a separate inventory forecast because the official system does not reflect operational reality quickly enough. These spreadsheets are not necessarily bad. Excel remains extremely useful. The problem appears when important business logic exists only inside unofficial files that are not controlled, documented or integrated with other systems. If AI is connected only to the official database, it may miss the information employees actually rely on to make decisions. Before introducing AI, management should therefore understand where critical information really lives.

Data Problems Often Begin With Process Problems

Businesses sometimes treat poor data quality as an IT problem, but the root cause frequently sits in the process that creates the data. If five employees can create customer records without checking whether the customer already exists, duplicates are predictable. If salespeople can submit orders without completing required customer information, missing fields are predictable. If different departments use different product codes, inconsistent reporting is predictable. Cleaning the database once may improve the situation temporarily, but the same problems will return unless the process changes. Management should therefore ask not only “How do we clean this data?” but also “Why did the data become dirty in the first place?” Sustainable improvement requires addressing both the existing records and the behaviour, controls and system design that continuously create new information.

Cleaning 100,000 Records Is Useless if Employees Create 1,000 New Problems Next Month

Imagine a company spends several months cleaning its customer database. Duplicates are merged, addresses standardised and missing information updated. Management celebrates because the system finally looks organised. Six months later, the same problems return because employees continue entering customers however they like. This is why data governance matters. The organisation needs practical rules around how important information is created, changed and maintained. Required fields may need to be defined. Duplicate checks may need to occur before new records are created. Naming conventions may need to be standardised. Responsibility for updating certain information should be clear. Data quality cannot depend entirely on occasional cleaning exercises. It needs to become part of normal operations.

One Customer ID Can Be More Valuable Than Another AI Tool

Businesses sometimes look for sophisticated technology while overlooking simple structural improvements. A unique customer identifier that is consistently used across accounting, CRM, billing and customer service systems can dramatically improve the organisation’s ability to connect information. The same principle applies to suppliers, products, employees and projects. Names are often unreliable identifiers because spelling, abbreviations and formatting change. Stable identifiers make it easier to reconcile information across systems and reduce ambiguity. This may not sound as impressive as deploying generative AI, but reliable identifiers create the foundation on which more advanced analysis depends. In many digital transformation projects, the boring structural decisions determine whether the exciting technology ultimately works.

Supplier Data Can Be Just as Messy as Customer Data

The same problem appears in procurement and accounts payable. One supplier might exist under several names because different employees created separate records. Bank information may be outdated, tax information incomplete and supplier categories inconsistent. Management then asks AI to identify opportunities for procurement savings. The analysis says the company spends S$300,000 with Supplier A and S$250,000 with Supplier B, when both records actually belong to the same supplier. The business therefore fails to recognise that it spends S$550,000 with one vendor and may have greater negotiating power than management realised. Poor supplier data can also make it harder to identify concentration risks, unusual payments or duplicate invoices. Better data therefore supports much more than AI. It strengthens ordinary financial and operational management.

Product Data Can Produce Expensive AI Recommendations

Imagine a retailer selling 10,000 products. Some products appear under old codes, new codes and regional codes. Product descriptions are inconsistent, costs are outdated and discontinued items remain active in the system. Management introduces AI to forecast demand and recommend purchasing quantities. If the system treats equivalent products as separate items or relies on outdated costs, the forecast may not reflect actual demand or profitability. The company could purchase too much inventory, fail to order products customers actually want or misunderstand which products generate attractive margins. Forecasting technology can be powerful, but its usefulness depends heavily on whether historical information is structured in a way that reflects the underlying business.

Historical Data Is Not Automatically Good Data

Companies often believe they are ready for AI because they possess ten years of transaction history. The quantity sounds impressive, but historical data may contain changes that make direct comparison difficult. Product codes may have changed, accounting systems may have been replaced, business units may have been reorganised and customer classifications may have evolved. A revenue category used in 2020 may not mean exactly the same thing in 2026. If these changes are not understood, AI may identify patterns that reflect changes in classification rather than genuine changes in the business. Historical depth is useful only when management understands what the data represents across time.

More Data Is Not Always Better

The assumption that AI becomes better whenever it receives more information can encourage businesses to collect everything without considering whether the information is useful. A company may have millions of records, but many may be outdated, duplicated or irrelevant to the question being asked. Feeding unnecessary information into an analysis can create complexity without improving decisions. Businesses should start with the problem they want to solve and identify the information required to solve it. If management wants to predict customer payment behaviour, for example, invoice dates, payment history, customer characteristics and credit terms may matter more than hundreds of unrelated fields. Data strategy should therefore be driven by business questions rather than by the belief that collecting more information automatically creates more intelligence.

Define the Question Before Building the AI Solution

A common mistake is starting with “We need AI” instead of “What problem are we trying to solve?” The first statement encourages the organisation to search for somewhere to use a fashionable technology. The second forces management to define an outcome. Perhaps the company wants to reduce customer response time, improve demand forecasting, identify overdue receivables earlier or reduce manual invoice processing. Once the objective is clear, management can identify what information the system needs and assess whether that information is reliable. This approach often reveals that different AI applications require very different levels of data readiness. A customer-service assistant may depend heavily on accurate product and policy information, while a forecasting tool may require several years of clean transactional data.

Not Every AI Project Requires Perfect Data

Businesses should also avoid going to the opposite extreme and delaying every AI initiative until every database is flawless. Perfect data is rarely realistic. The appropriate standard depends on what the AI system is being asked to do and what happens if it is wrong. A tool that helps employees draft internal meeting summaries may tolerate imperfections that would be unacceptable in a system recommending credit limits or producing information used for financial decisions. Companies should therefore prioritise data quality according to risk and business value. The objective is not to clean everything before doing anything. It is to ensure that the information supporting important decisions is sufficiently reliable for the intended purpose.

Start With the Data That Matters Most

If a company has hundreds of thousands of records, attempting to clean everything simultaneously can become expensive and overwhelming. A more practical approach is to identify critical data domains. Customer records may matter most for sales AI. Product and inventory data may matter for demand forecasting. Supplier and payment data may matter for procurement analysis. Financial data may matter for management reporting. By focusing on the information connected to a specific business objective, the company can make measurable progress without turning data quality into an endless transformation programme. Once governance improves in one area, the same principles can gradually be extended to others.

Decide Who Owns the Data

One reason data quality deteriorates is that everyone uses information but nobody owns it. Sales believes customer data belongs to IT because it sits inside the CRM. IT believes sales owns it because salespeople enter the information. Finance uses the customer record for invoicing but may not know who is responsible for updating addresses. When something is wrong, everyone can explain why another department should fix it. Businesses need clear ownership for important data. Ownership does not mean one person manually updates every record. It means someone is accountable for defining standards, monitoring quality and ensuring that problems are resolved. Without ownership, data-cleaning initiatives often produce temporary improvements before old habits return.

Finance Can Play an Important Role in Data Discipline

Finance teams are accustomed to reconciliations, supporting documents, consistency and controlled processes. Those habits can be valuable as businesses become more data-driven. Finance does not need to own every customer, product or operational dataset, but it can help establish a culture where important numbers can be traced and explained. If a management dashboard reports S$5 million of revenue, someone should understand how that figure connects to underlying transactions. If an AI system reports customer profitability, management should know which costs and revenue sources were included. The ability to reconcile information becomes increasingly important as businesses introduce more automated analysis.

AI Can Expose Problems That Were Already There

When an AI project produces inconsistent results, management may blame the technology. Sometimes the system is genuinely unsuitable. In other cases, AI simply reveals data problems the organisation has tolerated for years. Perhaps two departments have always used different customer classifications, but nobody noticed because their reports were reviewed separately. Once AI combines the information, the inconsistency becomes obvious. This can actually be useful. An AI implementation can force the company to confront questions about definitions, ownership and processes that should have been resolved earlier. The project may therefore create value even before the AI system produces sophisticated predictions, simply by exposing weaknesses in the organisation’s information foundation.

Different Departments Need the Same Definition of Important Numbers

Ask sales, finance and operations how many “active customers” the company has and you may receive three different answers. Sales might define active as anyone who purchased within 12 months. Finance may count customers with transactions during the financial year. Customer service may include anyone with an open account. None of these definitions is necessarily wrong, but they answer different questions. If management asks AI to analyse active customers without defining the term, the output may vary depending on which dataset is used. Businesses therefore need common definitions for important measures. Data consistency is not only about correcting spelling mistakes. It also means agreeing on what the information actually means.

The Same Problem Applies to Profitability

Management may ask an AI tool, “Which customers are most profitable?” Before the technology answers, the company needs to define profitability. Does it mean revenue minus direct product cost? Should delivery expenses be included? What about sales commissions, customer support time, discounts and returns? Different definitions can produce different rankings. AI can calculate the result quickly once the rules and data are available, but it cannot resolve every management judgement automatically. Companies need to understand the business logic behind the metrics they ask AI to produce.

Access Controls Matter When AI Can See More Information

Data readiness is not only about accuracy. It is also about who can access information. An AI assistant connected to multiple internal systems could potentially make information easier to retrieve than before. That convenience needs appropriate controls. An employee who previously had no access to payroll information should not necessarily gain access simply because an AI interface can search across databases. Businesses should consider permissions, confidentiality and data sensitivity when designing AI solutions. The objective is to make appropriate information easier to use without unintentionally weakening existing access restrictions.

Data Privacy Cannot Be an Afterthought

Companies may hold personal information about customers, employees and business contacts. Before sending information into AI tools, management should understand how the tool handles that data, what information is being shared and what internal policies apply. Employees should not assume that every convenient public AI service is an appropriate place for confidential company information. Businesses need clear guidance about approved tools and appropriate uses, particularly as employees increasingly experiment with AI independently. A strong AI strategy therefore involves governance as well as technology.

Human Review Still Matters

Even with clean data, AI outputs should be reviewed according to their importance and risk. An AI-generated marketing idea requires a different level of review from an AI recommendation affecting a customer’s credit limit or a significant financial decision. Human judgement remains important because business decisions involve context that may not exist in the dataset. A customer may appear financially weak based on historical numbers but have just received substantial new investment. A product may appear unprofitable but be strategically important because it drives sales of other products. Data informs judgement. It does not eliminate the need for judgement.

Do Not Automate a Decision You Do Not Understand

If management cannot explain how a decision is currently made, automating it can create new risks. Before asking AI to decide which customers should receive discounts, for example, the company should understand the factors its experienced employees currently consider. Customer size, payment history, strategic value, competitive conditions and margin may all matter. Mapping the existing decision process helps identify what data the AI system needs and where human judgement should remain. Automation is most effective when the company understands the process well enough to decide which parts should be automated and which should not.

Measure the Business Outcome, Not the Number of AI Tools

A company with ten AI subscriptions is not necessarily more advanced than a company with one well-designed application. Management should evaluate AI projects through outcomes such as reduced processing time, lower error rates, improved customer response, better forecasting or stronger conversion rates. Usage statistics can help, but they are not the ultimate objective. Employees generating 100,000 AI prompts does not prove that the company became more productive. Technology adoption should be connected to measurable business improvement. Otherwise, AI risks becoming another software category that increases costs without creating corresponding value.

Give the Project a Baseline Before You Start

If management expects AI to reduce invoice-processing time, measure the current processing time before implementation. If the goal is to improve customer response, establish the current response time and resolution rate. If AI is intended to improve forecasting, measure current forecast accuracy. Without a baseline, the company may implement the technology successfully but have no reliable way to determine whether anything improved. The same discipline applies to data quality. Measure duplicate rates, missing fields or reconciliation differences before cleaning so that progress can be demonstrated afterwards.

Good Data Creates Benefits Even if the AI Project Changes

One of the strongest arguments for improving data quality is that the benefits extend beyond any particular AI application. A clean customer database improves sales reporting, marketing, invoicing and customer service. Reliable supplier information improves procurement and payment controls. Consistent product data supports inventory management and profitability analysis. Even if the original AI project is cancelled or replaced, the organisation still benefits from better information. Data improvement should therefore not be viewed merely as an expense required to “feed the AI”. It is an investment in the company’s broader ability to understand and manage itself.

Your Existing Reports May Improve Before AI Does Anything

Once duplicate records are removed, definitions standardised and missing information addressed, management may discover that ordinary reporting becomes substantially more useful. Customer concentration becomes clearer, supplier spending can be consolidated correctly and product margins become easier to understand. Some questions management planned to solve with AI may suddenly be answerable using existing business intelligence tools or straightforward analysis. That is not a failure of the AI strategy. It means the company identified that its fundamental problem was information quality rather than lack of sophisticated technology.

Better Data Can Reduce Audit and Reporting Friction Too

Reliable data also supports financial reporting and assurance processes. When customer, supplier, inventory and transaction information is consistent, reconciliations can become easier and unusual differences may be identified earlier. Finance spends less time manually correcting records at month-end, while management receives information it can understand and defend. For a firm such as Lee & Hew Public Accounting Corporation, the wider lesson is relevant beyond AI: strong information processes support better accounting, clearer management reporting and more organised financial records. AI simply makes the quality of those foundations increasingly visible because businesses are asking technology to use their information in more sophisticated ways.

Do Not Turn Data Cleaning Into a Five-Year Project

Data quality can become intimidating because once companies start looking, they often discover problems everywhere. The answer is not necessarily a massive multi-year programme. Management can begin with one high-value use case, identify the information required, measure its quality and fix the most important weaknesses. A customer analytics project might begin by consolidating duplicate customer records and establishing common customer identifiers. A purchasing project might start by cleaning supplier records and standardising categories. Small, targeted improvements can create momentum while demonstrating practical business value. The goal is continuous improvement rather than waiting for a mythical day when every field in every system becomes perfect.

Build Data Quality Into Everyday Work

The most sustainable approach is to make reliable information a normal operational responsibility. Employees creating records should understand why required fields matter. Systems should prevent obvious duplicates where practical. Managers should review important exceptions. Data owners should monitor quality and resolve recurring issues. When processes change, reporting definitions should be updated. This approach is less dramatic than a one-time “data transformation”, but it is more likely to produce lasting results. AI increases the value of good information, but the responsibility for creating that information remains distributed throughout the organisation.

Conclusion: Before Asking Whether Your Company Is AI-Ready, Ask Whether Your Information Is Ready

Businesses should absolutely explore how AI can improve productivity, customer service, forecasting, financial analysis and decision-making. The technology is developing quickly, and companies that ignore useful applications may eventually find themselves operating less efficiently than competitors. However, adopting AI should not begin with the assumption that technology can compensate for years of inconsistent information practices. If customer records are duplicated, important fields are missing, different departments use conflicting definitions and critical business information lives inside unofficial spreadsheets, those weaknesses will follow the organisation into its AI projects. In some cases, automation can make the consequences more significant because unreliable information can now be processed and distributed much faster.

The encouraging part is that businesses do not need perfect data before they can begin. They need to understand the purpose of the AI project, identify which information matters to that purpose and improve the quality of that information to an appropriate level. A company planning customer analytics can start with customer identifiers and transaction history. A company exploring demand forecasting can focus on product, sales and inventory information. A business using AI for financial analysis can concentrate on the reliability and consistency of financial data. This targeted approach is more practical than trying to clean every record in the organisation before experimenting with anything.

Management should also remember that data quality is ultimately a business responsibility rather than an IT housekeeping exercise. Duplicates are created because processes allow them. Missing information exists because nobody required it. Conflicting definitions survive because departments were allowed to develop separate versions of the same metric. Unofficial spreadsheets become critical because formal systems do not provide what employees need. Fixing these problems therefore requires decisions about ownership, processes, controls and responsibilities, not simply better software.

The biggest opportunity may be that preparing for AI forces businesses to improve information they should have improved anyway. A company that knows exactly who its customers are, what they buy, how profitable they are, which suppliers it depends on and what its operations actually cost is already in a stronger position, regardless of which AI tools it eventually adopts. Clean, consistent and well-governed data improves ordinary management before it improves artificial intelligence.

So when the next AI proposal reaches the management meeting, the first question should not automatically be “Which AI platform should we buy?” A more useful starting question is “What information will this system rely on, and how confident are we that the information is actually correct?” If the company cannot answer that question yet, the AI project does not necessarily need to stop. But management has discovered where the real transformation needs to begin.