Skip to main content
Services

From messy source to
usable data

Thousands of photos without descriptions, an old archive, a product file full of gaps. We build AI tools that read, clean and enrich sources like these. Automated, verifiable and reusable.

Data generation

Manual work that does not have to be manual

Adding metadata to 3,000 photos. Describing products that only exist as an article number. Emptying an old system before it goes offline. Jobs like these always stayed on the shelf because they were too big for human hands. With AI they are suddenly feasible: we build the pipeline, you get clean data back in the format you need.

What we build

What we get out of your sources

  • Scraping and extraction

    Reading websites, archives and old systems automatically, even without an export or API.

  • Image recognition

    AI looks at your photos and writes descriptions, alt texts and keywords for them.

  • Metadata and EXIF/IPTC

    Writing descriptions straight into the files themselves, following the standards archives and photo systems read.

  • Text generation

    Product descriptions, summaries or SEO texts based on the data you already have.

  • Cleaning and normalising

    Duplicates removed, formats aligned, missing fields filled in.

  • Export to any format

    Excel, CSV, a database or straight into your CMS. The data lands where you work with it.

From practice

2,756 archive photos, described automatically

For a regional heritage project we built a data tool that reads an entire online image bank: 2,756 historical photos, including all their associated data. The tool removes watermarks, has AI write a description and alt text for each photo, stores that metadata in the files themselves (EXIF/IPTC) and delivers everything as an organised archive with an Excel overview.

Done by hand, this would have taken months. The tool runs incrementally, so new additions are picked up automatically. And because we calculated the AI costs per model up front, the client knew exactly where they stood.

Approach

How we approach a data project

  1. 1

    Examine the source

    We investigate what your source contains and what can be extracted. Often more than you think.

  2. 2

    Trial run on real data

    A small batch first, so you see the quality and we can fine-tune the instructions.

  3. 3

    Run the pipeline

    The full processing run, with logging and error handling. Large volumes run through the night.

  4. 4

    Deliver and repeat

    You get the data in your preferred format. The tool stays usable for future batches.

Frequently asked questions

Frequently asked questions about data generation

Good, but not flawless. That is why we always build in a review step: spot checks, a review screen or a score per item so doubtful cases rise to the top.

We calculate the AI costs per model in advance and pick the cheapest model that meets the quality bar. For thousands of images you are often looking at tens of euros in AI costs, not hundreds.

Not always. We check the terms and rights of the source beforehand. For your own data, licensed sources or public archives there is usually no obstacle.

Yes. We set up the pipeline as a scheduled task on your server, for example for new uploads or a daily sync with an external source.

Building smart solutions together?

Stijl