Anna versus the big AI: 4 of 14
If you build something slower than the competition, you should at least measure honestly how much slower. So I built Anna a test bench.
The experiment
There are three short German texts about an invented family: a profile, a family note and a weekly note. Anyone who has read them should know fourteen things, for example:
- Jonas lives in Leipzig and works at Kranich Logistik.
- He is married to Miriam, and Paul is Miriam's brother.
- Clara and Ben are Jonas' children. Clara is not Paul's daughter.
- Miriam gives Clara a book. Jonas no longer lives in Leipzig.
Two readers get the texts: Anna, on the iPhone, offline, with her hand-written German grammar, and a big AI, a language model in the cloud. A program counts what each of them writes into memory correctly. It also counts what they write in wrongly.
Anna: from 1 to 4
In the first measurement on 26 September, Anna scored 1 of 14, and she got that one point for a mistake she did not make. She answered many sentences with "Moin!" (a northern German "hi"), because her sentence classifier took them for a greeting.
Two days later she is at 4 of 14, and in no measurement so far has she written anything wrong. Where she gets stuck, she now says so:
"Clara ist die Tochter von Jonas." (Clara is Jonas' daughter.)
Anna: "I don't yet understand whose 'daughter' is meant within a statement. I haven't saved anything."
"Jonas arbeitet bei Kranich Logistik." (Jonas works at Kranich Logistik.)
Anna: "That's too many participants for me – I can't place 'Logistik' yet. Please tell me in two sentences, then I'll keep both."
Neither is a hit, but both are true. Anna doesn't guess, she says what she is missing, and the number of her follow-up questions has dropped from 15 to 8.
We trained on the test, by accident
I would rather not have written this part.
On 28 September, while double-checking, we noticed that two days earlier some training examples had gone into the app that looked suspiciously like the test sentences. A test you train on only measures itself. So we removed all of those examples and measured again.
Result: still 4 of 14, still 0 errors. The number was earned, the way there not entirely. Since then there is a fixed rule: the test bench is only for measuring, never for practising.
And the big AI?
The same three texts went to a large language model (Claude). The AI scored 12 of 14, clearly ahead, as expected.
More interesting is what it wrote into memory:
"Clara ist nicht die Tochter von Paul." (Clara is not Paul's daughter.)
AI: Clara – isNotDaughterOf – Paul
It invented a new relation, "is-not-daughter-of". That sounds right, but it isn't a negation.
"Jonas wohnt nicht mehr in Leipzig." (Jonas no longer lives in Leipzig.)
AI: Jonas – formerlyLivedIn – Leipzig
A clever reading. But because "Jonas lives in Leipzig" from the first text stays in memory, Jonas now lives there and no longer lives there at the same time.
"Miriam gibt Clara ein Buch." (Miriam gives Clara a book.)
AI: Miriam – gives – Buch and Miriam – givesTo – Clara
From these two loose statements you can no longer tell who gave what to whom.
| Hits | wrote something wrong | |
|---|---|---|
| big AI (26 Sep) | 12 of 14 | 4 invented relations, 1 contradiction, 1 duplicate person |
| Anna (28 Sep) | 4 of 14 | nothing |
To be honest, part of this is on us. fak10 doesn't just send the big AI the text, it also sends a protocol that tells it how to write facts down. When it was measured, that protocol had no field for "not". So it did what anyone in its place would have done and made up a word. We have revised the protocol several times since. The next episode shows what the AI makes of it.
What this is about
The big AI understands more, no doubt about it. Anna writes little today, but nothing wrong. We keep measuring both, in the open and with the same texts. So the question of this series is not whether Anna catches up with the big AI, but how far you can get without a cloud, without guessing and without a data centre, without ever writing anything wrong.
My goal for the launch is 10 of 14. The next episode comes when the number has moved.
— Maciek
If you want to be there when fak10 comes out: Join the waitlist