David Jansen van Vuuren

Product LeaderData & ML Platforms

I build and scale data and machine learning platforms. 18 years across e-commerce, adtech and fintech: search, recommendations, customer data platforms. Now building Routekeepers, a two-sided marketplace for permitted land access.

§ 01Selected impact

Numbers do the work.

Across roles at Takealot Group and Vicinity Media.
  • +20%
    search customer satisfaction

    Multi-intent ranking framework. Takealot.

  • 3×
    recommendation impression share

    Flexible context-aware API. Takealot.

  • Petabyte
    customer data platform

    Composable CDP, multi-brand, identity resolution. Takealot Group.

  • −15%
    zero-result searches

    Organic ranking model, Mr D groceries. Takealot Group.

  • 80% → 30%
    GeoIP-to-WiFi conversion target

    Micro Networks model. Vicinity Media (proof-of-concept).

  • 18
    years

    E-commerce, adtech, fintech. MBA. Cape Town.

§ 02ML Product Framework

A project execution framework for machine learning.

Built at Takealot Group to make the gap between exploration and productionisation explicit. The diamond below is the core visual: ideation expands into multiple candidate paths, then contracts into a single productionised model. Click any stage for detail.

The framework is about shape, not duration. Cadence varies with stack maturity, dataset readiness, and modern tooling (synthetic data generation, LLM-assisted prep, foundation-model baselines) often compresses several stages substantially.

EXPANSIONCONTRACTIONEXPLORATIONCOMMITUNSOLVEDCUSTOMERPAIN POINTIdeationRoadmapConsiderationRoadmapCommitmentDeliverable: Proof-of-conceptDeliverable: Customer-validated modelGeneratetraining dataCandidateModel 1CandidateModel nIterativeTrainingOffline POCEvaluationFinalise ModelArchitectureProductioniseStage ModelA/B Test ORInternal TestRetrainingModel MonitoringALLEVIATEDCUSTOMERPAIN POINTSTARTSHIPPED
Learnings — what changed after running the framework

Running the 6-month framework in practice surfaced two missing phases. The longer 15-month variant adds these as inline additions to the same diamond — not new structure, just acknowledgement that real ML work sometimes needs a dataset detour and a thinking pause.

  1. Learning 01
    Inserted before Generate Training Data

    Dataset Exploration Phase

    Often the bottleneck isn't the model — it's the dataset. Before committing to model training, generate multiple candidate datasets in parallel and pick the winner. Two evaluation paths: assess them via offline modelling against established metrics (when you'll need a model anyway), or — when the use case is simple enough — serve them directly as an A/B test. A caveat worth stating: if direct-serve datasets are sufficient on their own, you probably don't need a model at all.

  2. Learning 02
    Inserted after A/B Test, before retraining

    Investigation Buffer

    The A/B test doesn't always yield a clear answer. Reserve dedicated time after a test to investigate the result — understand the customer behaviour underneath the metric — instead of forcing immediate iteration. This stops cascading dependencies when an A/B test result requires deeper analysis (e.g. the Category Intent rollout, which needed investigation before promotion).

Both additions extend the project shape — a dataset arc up front, a thinking pause at the end. In practice they can add real time to a programme, though modern tooling (synthetic data generation, LLM-assisted analysis) compresses both phases substantially.

Adapting to agentic workflows

The diamond is calibrated for classical supervised machine learning — train a model against labelled data, evaluate offline, A/B test, productionise. The shape of the work still holds for agentic workflows: same expansion / contraction logic, same need for explicit checkpoints. Several stages just compress or change form.

Classical ML stage
Agentic equivalent
  • Generate training data
    Prompt design · RAG corpus · synthetic data
  • Candidate Model n
    Candidate prompts · tool sets · orchestration patterns
  • Iterative training
    Prompt iteration · retrieval tuning · tool-spec refinement
  • Offline POC evaluation
    LLM-as-judge · behavioural tests · red-team probes
  • A/B test or internal test
    A/B with completion rate · human eval · latency budgets
  • Productionise + stage
    Guardrails · rate limits · fallback prompts
  • Drift monitoring
    + hallucination rate · tool-call accuracy · latency

The two learnings (the dataset detour and the investigation buffer) carry over directly. The investigation buffer matters more in agentic work, not less — behavioural failures rarely have a clean numeric answer.

Companion artefacts

These are the artefacts that typically accompany a project on this framework. Names and roles only — the actual templates and specifications aren't included on this site; they live with the team that adopts the framework.

§ 03Experience

Career

MBA, University of Stellenbosch Business School (2013). BCom Management Sciences, University of Stellenbosch (2007).
  1. 2025 — Present

    Founder (Product)

    Building Routekeepers, a two-sided marketplace connecting verified riders to permitted land access and route guides. Designed and built the full product and data stack.

    MarketplaceData platformProduct
  2. 2024 — 2025

    Group Product Lead — Customer Data & ML

    Takealot Group

    Led product strategy for the customer data platform, CRM, insights and fraud ML across a multi-brand group, while keeping ownership of the ML roadmaps. Led a team of 4 product managers. Defined the architecture for petabyte-scale event collection, sessionisation and multi-model attribution. Established an extensible reporting framework and layered raw → aggregate → datamart conventions, improving data quality and coverage.

    Customer Data PlatformIdentity resolutionMulti-brand
  3. 2022 — 2024

    Group Product Lead — AI & ML

    Takealot Group

    Owned ML product strategy for search, recommendations, merchant ML and supply chain ML across Takealot, Superbalist and Mr D. Shipped organic ranking models at Superbalist and Mr D groceries, cutting zero-result searches (by 15% at Mr D) and lifting cart conversion.

    SearchRecommendationsML platforms
  4. 2020 — 2022

    Product Manager (AI/ML) — Search & Recommendations

    Takealot.com

    Owned the ML roadmap for search and recommendations, prioritised against the portfolio's North Star metrics. Led the multi-intent search framework: +20% customer satisfaction. Led the refactor of "You Might Also Like" (3x share of product page views, 5% to 15%) and launched a flexible recommendations platform for model-agnostic serving and faster experimentation.

    SearchRecommendations
  5. 2017 — 2020

    Product Lead — AdTech Portfolio

    Vicinity Media

    Owned strategy for a mobile location-advertising platform across three roadmaps: audience data, out-of-home attribution and ad server/campaign management. Led 2 product managers and defined success metrics for projects and releases. Developed out-of-home attribution linking billboard exposure to in-store visits. Founded the data science function, enabling ML-driven audience targeting and performance optimisation.

    AdtechAttributionAudience
  6. 2008 — 2015

    Consulting Manager

    Cape Value

    Led data analysis, market research and statistical modelling engagements, plus strategy and forecasting, for financial services (secured lending, claims management) and public sector (property revenue management) clients.

    ConsultingModeling
§ 05Contact

Get in touch