About PromptForge

A small team with unreasonable standards

PromptForge started in 2024 inside a marketing agency in Hamburg. We kept a shared doc of prompts that actually worked — the ones that survived contact with real client briefs, real deadlines and real code review. That doc grew to 400 entries, then colleagues asked for copies, then strangers did.

Today we're a fully remote team of five: two former agency operators, an ex-Google ML engineer, a technical writer, and a designer who thinks in Midjourney parameters. Between us we've run prompts against millions of tokens of production work — ad accounts, codebases, legal summaries, editorial calendars.

We sell prompt packs because we believe the gap between 'AI is amazing' and 'AI actually shipped this' is almost always prompt quality. Closing that gap is the whole company.

Methodology

Tested against real workloads, or it doesn't ship

Every prompt in every pack survives the same three-stage gauntlet. Roughly 40% of candidates fail. The ones that pass are the ones you buy.

01

Real workload trials

No toy examples. Marketing prompts run against live ad accounts and real landing pages. Coding prompts run against production repositories. If the output needs a full rewrite, the prompt fails.

02

Adversarial inputs

We feed each prompt vague briefs, missing context and contradictory instructions — because that's what real users paste. Prompts must ask clarifying questions or state assumptions instead of hallucinating confidence.

03

Blind cross-model scoring

Outputs are scored blind by the team across at least three frontier models. A prompt must beat the naive one-line version by a clear margin — otherwise you're paying for something you could have typed yourself.

Quality standards

What 'professionally engineered' means here

Versioned like software

Every prompt carries a version number and changelog. When a model update shifts behaviour, we re-test the whole pack.

Failure modes documented

Each prompt lists where it breaks: contexts it's bad for, inputs that degrade it, and the fix.

Structure over vibes

Prompts encode frameworks — positioning methods, review checklists, decision trees — not 'act as a rockstar' roleplay.

No filler

If a pack can be excellent at 85 prompts, we don't pad it to 200. Counts on the cards are the real counts.

The team

Five people, four time zones, one obsession

NK

Nadia Krüger

Founder · Prompt methodology

JM

Jonas Meireles

ML engineering · Eval systems

AO

Aiko Otanashi

Design · Midjourney pipelines

RB

Rory Byrne

Writing · Editorial packs

LS

Lena Stojanović

Operations · Customer success

See what survives our gauntlet

Browse the packs