An open invitation · co-authorship, not proofreading
Work on this with me.
I am not a classicist. I never took Greek or Latin, and I did not have the kind of education that hands you Thucydides with somebody standing next to you to say why it matters.
What I had was the wish to read these people directly, and the discovery that doing it is exhausting. A paragraph of Herodotus, then three search results to work out whether I had understood the paragraph, then back to the text having lost the thread of it.
So I built the thing I wanted to exist. Public-domain translations of 30 Greek and Latin works, with a note beside every passage supplying what that passage assumes you already know — who the person named without introduction is, what the coin was worth, which rule is being broken, where the joke is. And because readers differ, the amount of help is a dial: the same passage can be read bare, with a light gloss, or with the full apparatus open, and the reader chooses.
A language model wrote all 11,399 of those notes, and that is the part I cannot vouch for. I can build the machinery, verify the rights, pin every place on every map by hand, and write the automatic checks that catch the model inventing things — and I have. What I cannot tell you is whether the notes are any good. A note that is fluent, confident, plausible and wrong is precisely the kind of mistake I am not equipped to notice.
I built the dial as a courtesy to readers. It took somebody who does research for a living to point out what else it is: an instrument. A reading environment in which the assistance is a variable rather than a fixture — present, absent, prominent, buried, assigned rather than chosen if a design requires it — on texts that have held still for two thousand years. Nobody has had one before, and I do not have the training to use it.
That is what this page is for. Below are five questions, each addressed to a different field, each one I cannot answer on my own. If one of them is yours, I would much rather find out the answer with you than keep guessing at it.
Not a rubber stamp. I do not want a specialist to skim the site so that it can say "reviewed by a classicist" in the footer. That would make the site look better and teach nobody anything, and it is the one outcome this page exists to avoid.
Not free labour. The interesting version of this is a piece of research with a result, and a result has authors. Everything below assumes co-authorship and full access to everything I have — including the parts of the record that make the project look worse.
01
For classicists
How does machine commentary fail, and in what patterns?
The argument about automated scholarship is being conducted almost entirely without evidence, in both directions. What does not exist is a close description of how this kind of commentary actually fails when somebody qualified reads it: how often, on what sort of material, and in what recognisable patterns. That is a measurement problem with no instrument, and building the instrument is the harder half. It needs a scheme that separates an invented fact from a defensible reading pushed past its evidence, from an anachronism, from a claim that is perfectly true and useless — and, hardest of all, it has to handle the notes' claims about intent, tone and humour, where the notes assert the most and "wrong" is an argument rather than an error. Once the categories exist, the rates are arithmetic.
Intertextuality is the natural stress test, because an allusion is exactly the claim a fluent system can manufacture and a reader without the languages cannot check. The rule here is therefore absolute: the model may never assert that a passage quotes or alludes to anything, only explain a connection handed to it from a list a person compiled — which is why the site shows 1 quotations rather than several hundred. Lift that rule over passages whose intertexts are already well established in the scholarship, and you learn what share of genuine allusions the model finds and how much invention arrives with them — a measure of what the caution costs, and a finding that would inform every project of this kind.
The corpus makes the sampling tractable: 11,399 notes over 30 works spanning roughly a thousand years, every note held separately by kind, so author, genre and period can be varied deliberately rather than taken as they come. The result would be the first error taxonomy for machine commentary on classical texts, with measured rates attached.
02
For education researchers
Is the apparatus a scaffold, or a crutch?
The cognitive load literature says scaffolding helps novices. The same literature says the same support can hurt a reader one step further on — the expertise reversal effect — and that friction removed is sometimes learning removed: the looking-up I so resented doing may have been the desirable difficulty. Whether annotation in place is a scaffold or a crutch is exactly the kind of question that literature exists to settle, and as far as I can establish, nobody has asked it of readers meeting ancient texts in translation.
This is what the dial is for. The same passage can be served bare, lightly glossed, or fully annotated, and the level can be assigned rather than chosen. The outcomes worth measuring are the obvious three: comprehension now, retention at a delay, and the transfer test that actually matters — hand the reader a new passage with no notes at all and see what the apparatus taught them. If readers who had the full apparatus do worse on their own afterwards, I want to know that more than I want anything else on this page.
Building the conditions is straightforward and it is my job, not yours — along with consent screens, random assignment, and whatever else your protocol requires.
03
For psychologists
Does a supplied reading foreclose the reader's own?
Many of the notes take positions — on tone, register, intent, where the joke is — and the site flags them as arguments rather than facts. Flagging may not help. The question is whether a reader who meets the note before the passage ever really reads the passage, or whether the note's interpretation simply becomes theirs: anchoring, applied to literary judgement. The design is almost embarrassingly clean — the same interpretive note served before the passage, after the reader has committed to a reading of their own, or not at all, with agreement as the measure.
Its twin is calibration. Annotation in place may improve understanding, or it may mainly improve the feeling of understanding — readers whose confidence has grown faster than their comprehension. Metacomprehension is a measured thing, the confidence–accuracy gap is a standard instrument, and if this apparatus mostly manufactures conviction, that is a more consequential result than the one I am hoping for. It bears directly on how anything of this kind should be built from here, and it would still be worth publishing.
04
For translation studies
Can the machine catch a Victorian translator tidying?
Copyright drives much of this corpus into pre-1929 translations, which elevate register, moralise, and flatten obscenity and humour with some consistency. The notes make claims about precisely this, constantly — that the translator has softened a word, expanded a silence, or quietly declined to render something — and the Greek or Latin sits aligned beside every one of those claims, passage by passage.
It is the most tractable question on this page, because the claims have determinable answers: take the notes that accuse a translator of tidying, and check them against the original. It also has an audience well beyond this project. Reading Victorian translations is what most non-specialists actually do, and whether an automatic gloss can reliably flag where one has been tidied is worth establishing whatever one concludes about the rest of the method.
05
For HCI and science communication
What does the label "machine-written, unreviewed" actually do?
Every page on this site tells the reader, plainly, that the notes were written by a machine and reviewed by nobody. The assumption behind that honesty — mine, and increasingly everyone's — is that disclosure produces calibrated trust: told what the notes are, readers will lean on them accordingly. As far as I can find, nobody has tested whether readers of a working site see the label, believe it, or change a single behaviour because of it — how much they check, what they are willing to cite, how far they let a note settle a question.
The label is mine to vary: prominent or buried, repeated beside every note or stated once in a footer, reworded to any strength a design calls for, within a study readers have consented to join. Every archive, encyclopaedia and news organisation now attaching provenance labels to machine-written text is guessing at this. A measured answer would travel far beyond classics.
Smaller ways in
Nothing here needs a project, a grant or a meetingTwo hours, one work
Read a single text with the notes open and mark what is wrong, what claims too much, and what is simply missing. Rough and unsystematic, it would still be the first evidence anyone has.
One session, one seminar
Give a class a passage with the notes open and ask them to argue with the margin. An hour of students marking what a machine got wrong is pilot data, and I will build whatever version of the page the session needs.
Argue with a reading
Tone, register and intent are arguments rather than facts, and the site says so. A specialist who disagrees with a reading and says why is worth more to me than a correction — the corrections form takes those too.