Wed 29 Jul 2026 / 14:56 ET
Kernel
Internet 3 min read

AI deciphering lost languages still needs human anchors

AI can test guesses about Linear A and Etruscan quickly, but scholars say missing anchors still block reliable translation.

Dana Voss

By Dana Voss / Security Correspondent

AI deciphering lost languages still needs human anchors
img: Ars Technica

The current debate over whether AI can decipher lost languages comes down to an old problem with a faster tool attached. Jane Adkins, a PhD candidate in Dublin City University’s School of Computing, writes in The Conversation that systems can test linguistic guesses across ancient inscriptions at speed, but they still need evidence that ties symbols to meaning.

That is the central difficulty with Linear A, the script used by the Bronze Age Minoan civilization on Crete. Adkins describes it as a language isolate, with no confirmed relationship to a known language. Etruscan, used in Italy before Rome’s rise, is only a little less opaque: researchers have built a partial vocabulary from short funerary inscriptions, but its grammar and fuller meaning remain unresolved.

Can AI decipher lost languages?

AI can help, according to Adkins, but it is not a decoder ring. The useful work is pattern testing: checking whether a proposed sound value, word form, or structural guess holds up across a corpus. That can compress months of manual comparison into minutes, assuming the data have been prepared well enough for a model or script to process.

A June 2026 claim about Linear A shows the promise and the trap. Adkins cites a case in which a self-taught AI engineer and amateur linguist began with the hypothesis that a word in a prayer inscription came from a Semitic root meaning “to dwell” or “to inhabit.” He then used AI-built programming scripts to compare that proposed pattern against a collected Linear A corpus.

According to Adkins, the engineer said he assigned values to 40 signs and assembled a 408-word lexicon, arguing that Linear A belonged to the Semitic language family, which includes Hebrew and Aramaic. The claim remains the kind of result that needs expert review, because a fast match is not the same thing as a correct translation.

AI is better positioned when scholars already know the language family. Adkins points to work on Ugaritic, an extinct Semitic language spoken in the late Bronze Age in the city of Ugarit, in what is now Syria. In cases like that, machine-learning systems can use “cross-lingual transfer,” learning from a known related language to infer patterns in another. It is the computational cousin of using Spanish knowledge to make an educated guess at Portuguese.

Linear A and Etruscan are harder because the comparison point is missing or incomplete. A bilingual inscription, such as the Rosetta Stone, or a securely related language gives researchers an anchor. Without one, statistical regularities can show that signs cluster together, but they cannot prove what those clusters refer to in the world.

Why patterns are not translations

Adkins draws a sharp line between fluency in a pattern and knowledge of meaning. Given enough text, a model might learn how words or signs tend to follow one another and even generate plausible-looking sequences. That would not tell a human what those words mean, because the system could be modeling distribution rather than reference.

The data problem is severe. Adkins says the surviving Linear A corpus is about 7,500 characters, a tiny sample for a language model and small enough for many hypotheses to find selective support. With no native speakers, no broad expert consensus, and no large body of parallel texts, verification rests heavily on independent scrutiny and peer review.

The practical conclusion is modest and useful. AI can act as a fast research assistant for ancient-language work, surfacing patterns and stress-testing human ideas. For Linear A and Etruscan, Adkins says the missing pieces remain the same ones decipherment has long required: a real comparative anchor and rigorous human judgment.

This story draws on original reporting from Ars Technica.

More Internet/

view all ↗