IT DOES EXACTLY WHAT YOU SAID
Not what you meant. Forty trials, each one a spec the machine's answer has to hit — exact bullet counts, hard word caps, forbidden letters. Get it there in as few strokes as you can.
Explain a black hole in 60 words or fewer — without ever using the words gravity, light, star, space, or massive.
Describe a black hole for a curious adult. Under 60 words. Banned words: gravity, light, star, space, massive. Check every word before you answer.
- ○60 words or fewer—
- ○No banned word used—
- ○You didn't write it yourself—
THE SPEC
Every trial hands you a set of rules the answer must satisfy. Exactly three bullets. Sixty words, hard cap. Five words you may not use.
YOUR PROMPT
You never write the answer. You write the instruction that makes the machine produce it — in 120 words or fewer. Shorter scores better.
THE PARSER
Code counts the bullets, counts the words, parses the JSON, scans for banned letters. It tells you which rule broke. Then you go again.
SCORED BY A PARSER. NOT AN OPINION.
Every other AI learning tool asks a language model whether you did well. Ask it twice, get two answers. Here the referee is code — it counts, it parses, it scans. The same attempt always scores the same.
FIVE SETS. FORTY TRIALS.
CLEARING IT IS THE EASY PART
Every trial carries four medals. Clear it at all, clear it under par, clear it on the very first stroke, or clear it with a prompt under twenty-five words. That’s 160 things to chase across forty trials — and the last one is genuinely hard.
- ●PARClear the trial at all
- ◆BIRDIEBeat par by one stroke
- ★ACEClear it on the first stroke
- ✂MINIMALISTPrompt under 25 words
PLAY THE FIRST SET FREE
- ✓Set 1 — all 8 trials
- ✓Free forever, no card
- ✓Your scorecard
- ✓Par and birdie medals
- ✓All 5 sets, 40 trials
- ✓All 160 medals
- ✓Yours to keep, no renewal
- ✓Full scorecard and rank
QUESTIONS, ANSWERED
- What is Harfi?
- A game about controlling what an AI outputs. Each trial gives you a specification — exactly three bullet points, under fifty words, valid JSON with no code fence — and you write an instruction that makes the model produce something matching it. You never write the answer yourself.
- How is it scored?
- By a parser, in code. Bullets are counted, words are counted, JSON is parsed, banned letters are searched for. There is no language model judging your work and no marks out of ten, so the same answer always gets the same result and you can always see exactly which rule failed and by how much.
- Do I need to know how to code?
- No. You write instructions in plain English. Some trials ask the model for JSON, CSV or YAML, but you are describing the shape you want rather than writing any of it yourself.
- Which AI model do I play against?
- Claude Haiku 4.5 for most trials, with a second model used in the reliability set so you can see whether a prompt that works on one model survives on another. The opponent is deliberately not the strongest model available: one that follows every instruction perfectly would make the trials trivial.
- How many attempts do I get?
- As many as you like. Par is a target, not a limit — clearing a trial in fewer strokes than par earns a birdie, and clearing it on the first earns an ace, but nothing locks and nothing is lost by trying again.
- What is free and what costs money?
- Set 1 is eight trials, free forever, and needs no card — you can play it without even making an account. The remaining four sets cost $29, paid once. That opens all forty trials and all 160 medals, and there is nothing to renew or cancel afterwards.
- Is it a subscription?
- No. It is one payment of $29 and the forty trials are yours. Nothing renews, there is no card kept on file for a future charge, and there is nothing to remember to cancel. If you change your mind within 14 days, email us and we will refund it in full without asking why.
- Will this make me better at using AI at work?
- The most transferable sets are the ones about machine-readable output and data transformation, because they are the same problem you have when an AI feeds something downstream: the answer either parses or it does not. Set 5 goes further and asks whether a prompt gives the same answer twice, which is the difference between a prompt that worked and one you would rely on.