[{"data":1,"prerenderedAt":117},["ShallowReactive",2],{"\u002Fwriting\u002Fmanaging-product-data-at-scale":3,"\u002Fwriting\u002Fmanaging-product-data-at-scale-related":112},{"id":4,"title":5,"body":6,"description":89,"extension":90,"featured":91,"kind":92,"meta":93,"navigation":91,"ogImage":94,"order":95,"path":96,"publishedAt":97,"relatedWork":98,"seo":99,"stem":100,"summary":101,"takeaways":102,"technologies":107,"videoDuration":94,"videoId":94,"__hash__":111},"articles\u002Fwriting\u002Fmanaging-product-data-at-scale.md","How I Manage Product Data Across 200,000+ Ecommerce SKUs",{"type":7,"value":8,"toc":78},"minimark",[9,14,18,22,25,29,32,37,41,44,48,57,61],[10,11,13],"h2",{"id":12},"the-short-version","The short version",[15,16,17],"p",{},"Most of the pain in a large ecommerce catalogue isn't any single hard technical problem — it's the\naccumulation of a hundred small inconsistencies across suppliers, formats, and time. Here's the set\nof patterns I've found actually hold up at 200,000+ SKUs.",[10,19,21],{"id":20},"treat-every-source-as-untrusted","Treat every source as untrusted",[15,23,24],{},"It doesn't matter if a supplier feed has been reliable for two years — the day you stop validating\nit is the day a malformed export takes ten thousand SKUs to zero stock. Every ingestion job should\nassume the incoming file could be truncated, reordered, re-encoded, or just wrong, and validate\nbefore anything downstream trusts it.",[10,26,28],{"id":27},"matching-needs-a-confidence-score-not-a-boolean","Matching needs a confidence score, not a boolean",[15,30,31],{},"Product matching — deciding whether a new record is the same as an existing one — is rarely a clean\nyes\u002Fno. Exact key matches (a stable part number, a UPC) are easy. Everything else is a confidence\nproblem: title similarity, brand plus attribute overlap, price plausibility. The systems that work\nwell score that confidence and route anything below a threshold to a human, instead of forcing a\nbinary decision the data doesn't support.",[33,34],"architecture-diagram",{":steps":35,"title":36},"[{\"label\":\"Exact key match\",\"detail\":\"Auto-approved\"},{\"label\":\"High-confidence fuzzy match\",\"detail\":\"Auto-approved, logged\"},{\"label\":\"Low-confidence match\",\"detail\":\"Human review queue\"}]","Matching confidence routing",[10,38,40],{"id":39},"diff-before-you-publish","Diff before you publish",[15,42,43],{},"The single highest-leverage habit in catalogue automation is generating a diff report before\nanything goes live — new items, removed items, changed fields, all visible before publication. This\ncatches more real problems than any individual validation rule, because it surfaces things you\ndidn't think to write a rule for.",[10,45,47],{"id":46},"the-hardest-part-isnt-technical","The hardest part isn't technical",[15,49,50,51,56],{},"The genuinely hard part of running a catalogue this size isn't the code — it's the organizational\ndiscipline to keep supplier relationships, data contracts, and escalation paths clear enough that\nthe automation has something reliable to build on. I go deeper on the system itself in the\n",[52,53,55],"a",{"href":54},"\u002Fwork\u002Fcase-studies\u002Fcatalogue-automation","catalogue automation case study",".",[10,58,60],{"id":59},"related","Related",[62,63,64,71],"ul",{},[65,66,67,70],"li",{},[52,68,69],{"href":54},"High-Volume Ecommerce Catalogue Automation"," — the full case study",[65,72,73,77],{},[52,74,76],{"href":75},"\u002Fwork\u002Fprojects\u002Fcatalogue-pipeline-toolkit","Product Catalogue Pipeline Toolkit"," — the public, reusable version of these patterns",{"title":79,"searchDepth":80,"depth":80,"links":81},"",3,[82,84,85,86,87,88],{"id":12,"depth":83,"text":13},2,{"id":20,"depth":83,"text":21},{"id":27,"depth":83,"text":28},{"id":39,"depth":83,"text":40},{"id":46,"depth":83,"text":47},{"id":59,"depth":83,"text":60},"Practical patterns for ingesting, normalizing, matching, and validating a 200,000+ SKU ecommerce catalogue fed by FTP files, APIs, and scrapers.","md",true,"article",{},null,1,"\u002Fwriting\u002Fmanaging-product-data-at-scale","2026-04-20",[54,75],{"title":5,"description":89},"writing\u002Fmanaging-product-data-at-scale","The practical patterns I rely on to keep a large, multi-source catalogue clean, current, and trustworthy.",[103,104,105,106],"Treat every ingestion source as untrusted input, no matter how long you've used it","Matching needs a confidence score and a human review queue, not a hard yes\u002Fno","Diff-before-publish catches more real problems than any single validation rule","The hardest part of catalogue automation is organizational, not technical",[108,109,110],"Python","PostgreSQL","Pandas","RifS8phw0wNELYSMrDwaHv_oCOBlz6BAps1Yg3GTPJ8",[113,115],{"path":54,"title":69,"summary":114},"Automating ingestion, cleanup, matching, and pricing for 200,000+ supplier SKUs.",{"path":75,"title":76,"summary":116},"A reusable Python toolkit for importing, normalizing, validating, and comparing supplier catalogue files.",1785688079476]