Christian M. Andersen · senior backend developer, Copenhagen archive · 2026 · 5 channels open

cma:~$ cat words/going-smaller.md

Recovered record ·· words.0008

Going Smaller

This article is a follow-up to Did You Mean?, about what happened when I stopped picking models off the shelf and trained my own.

Last time, the guesser ended up as three signals working together: string comparison for typos, word counting for names, and a small language model for meaning. It worked. But every decision it made came down to a pile of thresholds I had set by hand, by squinting at 30 test URLs. Redirect if the word score is above 0.6, and at least 0.3 ahead of the next page, unless… You get the idea.

That bothered me. Not because it was wrong, but because I had no way of knowing how wrong it was.

80 websites that don’t exist

To train anything, you need examples. Lots of them. Broken URLs, and the page each one was actually after. I don’t have a year of 404 logs lying around, so I had them written instead: a team of AI agents, each with a different brief. One played the web server reading its own 404 log. One was an SEO consultant who has migrated a hundred sites. One was an adversarial tester whose only job was to make the guesser redirect somewhere stupid.

The trick is that almost none of it is about this site. It’s 80 invented websites: a bakery, a Danish football club, a COBOL programmer’s memoir, a tattoo studio. Because if you train a model on cma.dk, it learns cma.dk. Train it on 80 other sites, and it has to learn the actual skill: what a person typing /that-weather-thing probably meant.

And then the important bit. Ten more invented sites, plus a hand-checked set of cma.dk examples, went into a frozen test set. Written before any training happened, never trained on, never tuned against. Every number in this post comes from there.

Bigger is better, right? (Again.)

You’d think I learned my lesson last time. I tried bigger models anyway. Gemma and bge were bigger and slower without being better. A multilingual model was genuinely better at Danish URLs, but the site is in English, and it was two and a half times the size.

What did help was going the other way. A static model is basically a dictionary: every word maps to a list of numbers, and a sentence is the average of its words. No neural network at all. I took one of those and fine-tuned it on the training data, which took seconds. Then I shrank it to half its size, and it scored exactly the same. Not roughly. Exactly.

A dozen numbers

The thresholds got replaced too. Instead of me guessing the cut-offs, a tiny model learns them: a logistic regression. It looks at all three signals for every page on the site and outputs one number, the probability that this is the page you meant. The whole thing is ten weights and a couple of thresholds in a JSON file. The biggest ones:

lexical_gap     1.59
semantic_gap    1.41
whole_gap       0.96
whole           0.88
semantic        0.26

What it learned surprised me. The biggest weights aren’t on how well a page matches. They’re on how much better it matches than every other page. A strong match that two pages share is worth almost nothing. A modest match that only one page has is worth a lot. Which, in hindsight, is exactly what all my hand-tuned “lead” thresholds were fumbling towards.

So, did it work?

On this site: wrong redirects went from 10 to 4 on the test set. That’s the number I care about most, because landing on the wrong page is worse than a 404 with good suggestions.

It’s also more careful. /weather used to send you straight to the DMI project. Now it shows it as a suggestion, because “weather” alone just isn’t that much to go on. You win some, you lose some.

And it got cheaper. The language model is gone, and so is everything it needed to run. A guess now takes about 0.03 milliseconds, twenty times faster than before, the server image shrank by 40%, and it uses less memory too.

My favourite bug from the whole thing: /wordpress didn’t work. Not because the model got it wrong, but because the spam filter in front of it treated anything containing “wordpress” as a bot probing for login pages. On a site with a post about WordPress. Oops.

Final thoughts

It’s tempting to think “training your own model” means going bigger. Heck, that’s what I assumed going in! But what came out the other end is smaller than what went in, and it makes better decisions with a dozen numbers than I did with a dozen thresholds.

The real test is still to come, though. Every number in this post was measured against examples I had written. The day this site has a few months of real 404s, I’ll find out how well it actually guesses.

‹ back to words