LLMs View History Through Today’s Lens

Large language models reconstruct history but view it through modern assumptions, introducing stereotypes and anachronisms. Researchers build time-locked models like Talkie and TimeCapsuleLLM, create temporal benchmarks, and test unlearning techniques to improve accuracy. Yet contamination and data limits persist. Real progress requires better grounding in period sources.
LLMs View History Through Today’s Lens
Written by Ava Callegari

Frontier AI systems promise vast knowledge. Ask one about 19th-century Ireland and it might spin a tale in convincing brogue. The words flow. The facts line up. Yet something feels off. The voice carries a modern echo. Stereotypes creep in. An object appears that no one in that era would recognize.

Researchers have a phrase for it. “A history of the past isn’t the past.” So said Matthew Wilkens, associate professor of information science at Cornell University, in Communications of the ACM. The article, published July 27, 2026, lays bare a stubborn flaw. Large language models reconstruct events. They do not inhabit them.

Training data stretches across centuries. Models absorb patterns from billions of tokens. The result looks complete. But context slips away. Speech patterns evolve. Daily objects change. Values shift. LLMs flatten all of it into a confident, homogenized narrative. Often that narrative tilts toward a 21st-century Western perspective.

Ted Underwood, professor of information sciences and English at the University of Illinois, Urbana-Champaign, calls the training process a homogenization machine. “There’s a homogenization process that takes place during training, so these models are not able to speak authentically from a historical perspective,” he told the ACM publication. What emerges can sound authentic at first. Then it reveals itself as caricature. Underwood labeled one example “leprechaun talk.”

Modern Assumptions Distort the Record

David Bamman, associate professor at the University of California, Berkeley School of Information, points to a deeper uncertainty. Frontier models nail major historical facts. The trouble lies in verification. “We can’t be sure whether we’re measuring a model’s true reasoning ability or its ability to retrieve an answer someone else has already given,” Bamman said in the same ACM report.

Hamed Yaghoobian, assistant professor of computer science at Muhlenberg College, put it another way. Prompt a model to act Victorian. It performs like a costume drama. “It cannot unlearn the 20th century.”

These observations come at a moment when new research sharpens the critique. A NeurIPS 2025 poster introduced TimE, a multi-level benchmark that tests LLMs on real-world temporal challenges. It highlights three obstacles long ignored: dense temporal information, fast-changing event dynamics, and complex dependencies in social interactions. Experiments on both reasoning and non-reasoning models revealed persistent gaps. Performance varies sharply across scenarios. Test-time scaling helps in some cases. It does not close the blind spots. (NeurIPS, https://neurips.cc/virtual/2025/poster/121417)

Another 2025 paper in ACL Findings tackled explainable temporal reasoning. Authors found that LLMs often fail to produce convincing explanations when limited to text alone. Hallucinations appear. Structured knowledge from temporal knowledge graphs improves results. The work proposes GETER, a structure-aware framework that bridges graphs and text. Even tuned models stumble on cooperative signals in event chains. GETER identifies them more reliably. (ACL Anthology, https://aclanthology.org/2025.findings-acl.378.pdf)

Video models face parallel struggles. TemporalVLM, detailed in recent work from Retrocausal, addresses shortcomings in long-video understanding. Earlier systems treat clips as single units. They pool features or aggregate queries without time sensitivity. The new approach delivers finer timestamp prediction and caption accuracy. Previous models hallucinate objects or repeat descriptions. (Retrocausal, https://retrocausal.ai/video-llms-for-temporal-reasoning-in-long-videos/)

Yet the core historical problem remains. Kaspar Beelen, technical lead at the University of London’s School of Advanced Study, cut to it. “Data matters more than architecture.” More letters, posters, ledgers, and newspapers could help. Digitization costs time and labor. Censorship and human bias persist anyway.

Data contamination adds another layer. Even models trained with strict cutoffs absorb later knowledge. Footnotes slip in. Metadata leaks. Library catalogs mention future events.

The Talkie project illustrates the issue perfectly. Developed by Nick Levine, David Duvenaud, and Alec Radford, Talkie-1930-13B trained on 260 billion tokens of pre-1931 English text. Books, periodicals, court records, patents. Cutoff: December 31, 1930. The team built an anachronism classifier to filter the corpus.

It was not enough. The 13-billion-parameter model still knew details of the Roosevelt presidency, the New Deal, World War II, the United Nations, and the division of Germany. “An earlier 7B version of talkie clearly knew about the Roosevelt presidency and New Deal,” the project site notes. “talkie-1930-13b is additionally aware of some details related to World War II and the immediate postwar order.” The developers plan to scale to GPT-3 level this summer and push the corpus past a trillion tokens. (Talkie LM, https://talkie-lm.com/introducing-talkie)

A parallel effort started smaller. Hayk Grigorian, then a student at Muhlenberg College working with Yaghoobian, created TimeCapsuleLLM. Trained from scratch on roughly 90 gigabytes of London texts from 1800 to 1875. Books, periodicals, legal documents. No modern fine-tuning. The model does not pretend to be old. It simply reflects its data.

One test prompt began, “It was the year of our Lord 1834.” The output referenced real protests and Lord Palmerston’s actions. Grigorian had to Google to confirm the events actually happened. The model had inferred historical patterns from thousands of documents without explicit labels. Coverage in Ars Technica from August 2025 captured the surprise. “I was interested to see if a protest had actually occurred in 1834 London and it really did happen,” Grigorian wrote on Reddit. The project now spans multiple versions with over 136,000 documents. (GitHub, https://github.com/haykgrigo3/TimeCapsuleLLM; Ars Technica, https://arstechnica.com/information-technology/2025/08/ai-built-from-1800s-texts-surprises-creator-by-mentioning-real-1834-london-protests/)

These experiments show promise. They also expose limits. Wilkens explores AI simulations of past events. Replay moments. Generate scenarios. Compare against the historical record. Underwood’s group built a benchmark with 714 questions drawn from period sources. Books, newspapers, diaries, religious texts, archives. Questions test knowledge and inference across “parallax” categories. Answers shift depending on vantage point. A British colonial administrator sees events one way. A local citizen sees them differently.

Bamman pursues multimodal paths. His team digitizes films from 1922 onward. They track the rise of cinematic techniques. Close-ups. Camera movements. The work blends computer vision, natural language processing, and AI to illuminate cultural history. One paper appeared in PNAS. Another in EACL 2026. (PNAS, via ACM reference)

Machine unlearning offers another route. Strip out facts that postdate the target era. A 2025 arXiv paper explores anachronism removal. The hope is cleaner temporal grounding. (arXiv, referenced in ACM)

Progress accumulates. Yet researchers remain measured. “A lot has been lost and there are limits to what we can reconstruct,” Underwood said. “But we are making progress.” Wilkens agrees the goal is not simulation of conversations with Washington or Genghis Khan. It is historical grounding. Better understanding of the past. Clearer lessons for the present.

LLMs will keep improving. New benchmarks will test them harder. Time-locked models will scale. Multimodal data will add texture. Contamination will prove stubborn. The modern looking glass may never fully shatter.

But the effort matters. Historians gain tools. Educators gain nuance. The public gains context that goes beyond confident recitation. The past stays complicated. So do the machines we ask to explain it.

Subscribe for Updates

GenAIPro Newsletter

News, updates and trends in generative AI for the Tech and AI leaders and architects.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us