Product matching engine for medical supply data
A UK data-quality software vendor
An on-premise .NET engine that matches and cleans around 6 million medical product records from public and supplier sources, and explains every match it makes.
The problem
The vendor's first customer combines product data from the FDA, NHS Supply Chain, the MHRA, manufacturers and barcode scans: about 6 million products. Matching them meant hand-written SQL and a lot of manual fixing.
I owned it from requirements to handover: discovery with the vendor and their customer, a product requirements document, the engine, the review screens and the documentation.
- 6 million
- product records
- 4
- independently testable layers
What I built
- A .NET library that runs inside the vendor's Windows desktop product, with no server or extra install for the customer
- Barcode validation, including GS1 check digits, before any record enters the product master
- Extraction of strength, size and pack size from free-text descriptions, so records are compared field by field
- Layered matching: exact barcodes first, then rules, fuzzy matching and text similarity
- Every match explained attribute by attribute, with low-confidence pairs sent to a review queue instead of merged
A signed-off requirements document with measurable success criteria before any engine code, and a handover pack so the next developer is productive in a day.
Tools: C#, .NET, ML.NET, SQLite, SQL Server, FuzzySharp, CsvHelper, xUnit
