[{"data":1,"prerenderedAt":127},["ShallowReactive",2],{"\u002Fwork\u002Fprojects\u002Fcatalogue-pipeline-toolkit":3},{"id":4,"title":5,"body":6,"demo":102,"description":103,"extension":104,"featured":105,"github":102,"meta":106,"navigation":105,"ogImage":102,"order":107,"path":108,"plannedImprovements":109,"problem":114,"publishedAt":115,"seo":116,"status":117,"stem":118,"summary":119,"technologies":120,"__hash__":126},"projects\u002Fwork\u002Fprojects\u002Fcatalogue-pipeline-toolkit.md","Product Catalogue Pipeline Toolkit",{"type":7,"value":8,"toc":90},"minimark",[9,14,24,28,31,35,40,44,47,51,63,67,70,74],[10,11,13],"h2",{"id":12},"one-sentence-explanation","One-sentence explanation",[15,16,17,18,23],"p",{},"A command-line Python toolkit that turns messy supplier catalogue exports into normalized,\nvalidated, diffable data — the reusable core of the pipeline described in the\n",[19,20,22],"a",{"href":21},"\u002Fwork\u002Fcase-studies\u002Fcatalogue-automation","catalogue automation case study",", rebuilt from scratch\nwith synthetic data so the design and code can be public even though that case study's data isn't.",[10,25,27],{"id":26},"the-problem-being-solved","The problem being solved",[15,29,30],{},"Supplier catalogue ingestion is one of the most common — and most repeatedly reinvented — problems\nin ecommerce tooling. Every business writes its own one-off scripts to parse CSV\u002FExcel exports,\nguess at unit conversions, and check for obviously broken rows before publishing. None of that logic\nis inherently proprietary; it's just rarely extracted into something reusable.",[10,32,34],{"id":33},"planned-architecture","Planned architecture",[36,37],"architecture-diagram",{":steps":38,"title":39},"[{\"label\":\"Adapter\",\"detail\":\"Parse CSV\u002FExcel\u002FJSON into a common schema\"},{\"label\":\"Normalize\",\"detail\":\"Units, categories, attribute mapping\"},{\"label\":\"Validate\",\"detail\":\"Rule-based checks, typed errors\"},{\"label\":\"Diff \u002F export\",\"detail\":\"Compare snapshots, write clean output\"}]","Toolkit — planned shape",[10,41,43],{"id":42},"technology-stack","Technology stack",[15,45,46],{},"Python, with Pydantic for schema validation and typed records, pandas for the bulk transform steps,\nand a small CLI (Typer or argparse) so the toolkit is usable standalone or importable as a library in\na larger pipeline.",[10,48,50],{"id":49},"current-status","Current status",[15,52,53,57,58,62],{},[54,55,56],"strong",{},"Concept — design complete, implementation not yet started."," This page documents the toolkit I\nintend to build by generalizing the ingestion and validation logic from my professional work into a\nstandalone, MIT-licensed package with synthetic sample data. No public repository exists yet — check\nback, or ",[19,59,61],{"href":60},"\u002Fcontact","get in touch"," if you'd like to see this prioritized.",[10,64,66],{"id":65},"what-i-expect-to-learn","What I expect to learn",[15,68,69],{},"Extracting private pipeline logic into a general-purpose tool forces real API design decisions that\na one-off internal script never has to make — this is where I expect the most useful lessons about\nwhat's actually reusable versus what's specific to one business's data quirks.",[10,71,73],{"id":72},"planned-improvements","Planned improvements",[75,76,77,81,84,87],"ul",{},[78,79,80],"li",{},"Pluggable adapters for common supplier feed formats",[78,82,83],{},"Configurable normalization rules per source",[78,85,86],{},"Human-readable diff reports between catalogue snapshots",[78,88,89],{},"A tested validation-rules layer, not inline script logic",{"title":91,"searchDepth":92,"depth":92,"links":93},"",3,[94,96,97,98,99,100,101],{"id":12,"depth":95,"text":13},2,{"id":26,"depth":95,"text":27},{"id":33,"depth":95,"text":34},{"id":42,"depth":95,"text":43},{"id":49,"depth":95,"text":50},{"id":65,"depth":95,"text":66},{"id":72,"depth":95,"text":73},null,"A planned open-source Python toolkit that extracts the reusable parts of supplier catalogue ingestion — import, normalize, validate, diff — away from any one employer's private pipeline.","md",true,{},1,"\u002Fwork\u002Fprojects\u002Fcatalogue-pipeline-toolkit",[110,111,112,113],"Pluggable adapters for common supplier feed formats (flat file, JSON, XML)","A rules-based normalization layer (units, categories, attributes) configurable per source","A diff command that outputs a human-readable change report between two catalogue snapshots","Validation rules as a first-class, testable component rather than inline script logic","Every ecommerce business with a supplier-fed catalogue rebuilds the same handful of primitives — parse a messy CSV\u002FExcel export, normalize units and categories, validate before publishing, and diff one version of a catalogue against the next. That logic usually lives locked inside one company's private pipeline instead of existing as a reusable tool.\n","2026-04-01",{"title":5,"description":103},"concept","work\u002Fprojects\u002Fcatalogue-pipeline-toolkit","A reusable Python toolkit for importing, normalizing, validating, and comparing supplier catalogue files.",[121,122,123,124,125],"Python","Pandas","Pydantic","CSV \u002F Excel \u002F JSON","CLI","Hri19A0-m6yzVe7VaOMWBxJqoQgWK1BUMS2nv80D1W8",1785688079658]