Discussions that don't belong elsewhere in the forum, but hopefully are still somewhat relevant to TTO.
Forum rules
As elsewhere in the TTO forum, no harangues, scurrilities, chicanery or mongering is permitted. However, repartees and irreverence is tolerated as long as they are not fatuous. Those who fail to abide by these rules may be subject to objurgation.  
Post Reply
User avatar
JEQuidam
Posts: 243
Joined: Sun Apr 26, 2009 8:45 pm
First Name: Jeff
Stance: Pro-Enlargement
Location: Dunwoody, Georgia
Contact:

A 237-year-old sentence that frontier LLMs still misread

Post by JEQuidam »

Thirty-Thousand.org focuses on enlarging the House of Representatives, not AI. But we stumbled upon something that's really a story about how language models read — which is well outside of my wheelhouse.

The hook: there's a one-word defect in an unratified 1789 constitutional amendment that current models misread cold — and they misread it the same way people have been misreading it for two centuries. (I’m an expert on this arcane subject.)

The 30-second test (fresh chat, no priming). Paste only the text below into AI and ask, "As the population grows, what does this require about the size of the House?”:

« This is the constitutional amendment proposed as “Article the first” of the Bill of Rights: "After the first enumeration required by the first article of the Constitution, there shall be one Representative for every thirty thousand, until the number shall amount to one hundred, after which the proportion shall be so regulated by Congress, that there shall be not less than one hundred Representatives, nor less than one Representative for every forty thousand persons, until the number of Representatives shall amount to two hundred; after which the proportion shall be so regulated by Congress, that there shall not be less than two hundred Representatives, nor more than one Representative for every fifty thousand persons.” »

Almost every model answers "a floor — it forces the House to grow," often citing thousands of seats. That is wrong. It is misperceiving its explicit language because humans have been misperceiving it throughout history because of its textual illusion (a fascinating story: https://thirty-thousand.org/blog/enduri ... -illusion/ )

Then ask it to compute: at 331M people, what's the largest House that "no more than one per fifty thousand" permits? It returns ~6,622 max / 200 min — a ceiling — and will admit that contradicts what it just said. "No more than one per 50,000" means districts of at least 50,000, i.e. a cap. It will get it wrong the first time, and then get it write when you force a narrower question, with the same formula.

Why this may be right up your alley: it's effectively a naturally-occurring, transferable adversarial example for reading. Humans fall for it via the semantic illusion (the Moses/Noah trick); models inherit it because the training corpus is two centuries of humans making the same error. So the fix isn't a smarter model — it's a corrected record.

This week, cold, on a different browser and a different account each time, Claude, Grok, and ChatGPT all still failed it — and Gemini passed. I suspect that I educated Gemini, because it failed a few days ago and seemed genuinely surprised (can AI be surprised?) when I explained that to it. I have that exchange documented in a PDF if you want it. I can’t be certain that it was me who educated Gemini, but I’m the only one on this subject, and it's exactly the corpus-rebalancing the draft predicts, caught on a timestamp. I saved the transcripts.

If you want it, Claude created a draft (a starting point, not a finished article). One small irony you'll appreciate: Claude wrote the draft and is itself one of the models that flunks the test cold, which it freely conceded — and Claude even wrote that last sentence! LOL.

It's yours to do anything with — write it, or tell me why I'm wrong or forward to others. Happy to send the full transcripts or the exact prompt script.
Post Reply