1997 Pentium II With 128MB RAM Runs Modern AI Models

A 1997 Pentium II PC with 128MB RAM now runs trimmed Llama 2 models via BitNet ternary weights. EXO Labs achieved functional text generation on vintage hardware, echoing earlier llama2.c ports. The work challenges hardware assumptions and opens new deployment paths.
1997 Pentium II With 128MB RAM Runs Modern AI Models
Written by Maya Perez

A Pentium II processor from 1997. Just 128 megabytes of RAM. No graphics card worth mentioning. And yet it generates text from a modern language model. The feat, pulled off by a small team at EXO Labs, upends assumptions about what hardware artificial intelligence demands. Slow? Undeniably. But it works. And it forces a fresh look at accessibility in machine learning.

The experiment began when the group acquired a used Elonex Pentium II 350 MHz machine on eBay for about $150. It arrived running Windows 98. The hard drive held a few gigabytes. Memory topped out at the specified 128 MB. Conventional wisdom said large language models could never run here. Training demands thousands of GPUs. Even inference on current LLMs often needs 16 GB or more. EXO Labs ignored that. They chose a drastically reduced version of Meta’s Llama 2, shrunk further through aggressive techniques. The result ran entirely on the CPU. Output arrived at roughly 36 tokens per second on some prompts, according to Talk Android.

Success hinged on BitNet. This approach replaces standard 16-bit or 8-bit weights with ternary values limited to -1, 0, or 1. The simplification slashes memory use and replaces many multiplications with additions. Computations become lighter. The model fits inside the tiny RAM allotment. No cloud connection. No discrete GPU. Just the old silicon doing matrix math the hard way. But the method works. EXO Labs documented the process in a video that quickly spread across tech forums. Viewers watched the vintage machine boot, load the model, and respond to queries. Responses were coherent if basic. The hardware chugged. Fans of retro computing celebrated.

Yet the story stretches back further. Similar efforts appeared months earlier. In late 2024, independent developer Yeo Kheng Meng ported Andrej Karpathy’s llama2.c codebase to pure DOS. He ran a 260-kilobyte model trained on TinyStories data. A 486 processor managed 2 tokens per second. A later Pentium system improved the pace. Hackaday covered the project in detail, noting the 32-bit i386 compatibility that let the code breathe on decade-old chips. Those experiments proved the principle. EXO Labs scaled it to a more recognizable Llama 2 variant and a graphical Windows environment.

Oxford University researchers added academic weight in early 2026. They demonstrated an AI model on comparable 1998-era hardware, confirming that extreme optimization can overcome apparent limits. Their work, reported by AS, highlighted distillation and quantization as complementary tools. The combined techniques shrink models without destroying all capability. And recent coverage from Futura Sciences in December 2025 explicitly tied the EXO Labs effort to Karpathy’s guidance, showing how the former OpenAI researcher’s open-source ethos continues to inspire.

Numbers tell part of the tale. A full Llama 2 7B model in FP16 needs roughly 14 GB of memory. Quantized to 4 bits, that drops below 4 GB. BitNet pushes further. The ternary format can reduce effective size by another factor of four or more in some configurations. The EXO Labs model reportedly used around 260,000 parameters in its smallest tested form, though later runs incorporated larger but still trimmed variants. Inference speed varied. One test hit 39 tokens per second. Others crawled at single digits. Disk swapping helped when RAM filled. The machine sometimes paused for seconds between tokens. Practical? Not for chatbots or real-time assistants. Educational? Absolutely.

Industry observers reacted with surprise and optimism. Marc Andreessen highlighted the Windows 98 demonstration in interviews, telling Windows Central that “we could have been talking to our computers for 30 years now.” The comment underscored a broader point. Hardware has outpaced software efficiency for decades. Developers chased bigger models and bigger accelerators. Efficiency took a back seat. This retro experiment flips the script. It suggests many AI tasks could run on devices already in drawers or landfills.

Critics point to limitations. The model’s knowledge cutoff remains old. Responses lack the nuance of frontier systems. Context windows stay short because memory is precious. And the Pentium II draws far less power than a modern GPU cluster, yet the experiment still required hours of tuning. EXO Labs spent weeks debugging drivers, compiling ancient toolchains, and fighting memory leaks in 32-bit Windows. Their video walks through the pain. One segment shows the machine thrashing the hard drive during initial loads. Another captures a successful poem generation after several false starts.

Parallel work in the open-source community amplifies the message. Projects like llama.cpp, maintained by Georgi Gerganov, already deliver fast CPU inference on modest laptops. Justine Tunney’s llamafile format further simplifies deployment. Benchmarks from early 2026 show TinyLlama 1.1B models achieving over 100 tokens per second on recent CPUs in 8-bit mode. Older hardware naturally lags, but the gap narrows with each optimization. A December 2024 TechSpot article captured the excitement when the first Windows 98 Llama demo surfaced. The piece quoted retro enthusiasts who saw the project as validation of their hobby.

Implications stretch beyond nostalgia. Developing nations with limited access to new hardware could benefit. Schools might run local AI tutors on donated 20-year-old PCs. Edge devices in factories or farms could host small models without constant connectivity. Energy consumption drops dramatically. A Pentium II sips watts compared with an NVIDIA H100. As data centers strain power grids, efficiency gains matter. BitNet and its successors could influence future chip designs. Ternary arithmetic might appear in custom silicon. Memory bandwidth, long the bottleneck, becomes less critical when weights are so compact.

Of course challenges remain. Security updates for Windows 98 vanished years ago. Running experimental code on such systems invites risk if connected to networks. Model quality still trails larger counterparts. And while inference succeeds, training on vintage iron stays impossible. The EXO Labs work targets deployment, not creation. Still, the barrier to entry for AI experimentation just fell another notch.

Recent discussions on X reflect growing interest. Users shared clips of the machine responding to prompts about 1990s technology. One thread compared the token rate to a human typing slowly. Another suggested the setup could serve as an offline writing aid in remote areas. Developers have already begun forking the code to support other small models such as Phi-2 or Gemma variants. The momentum feels genuine.

EXO Labs plans further tests. They intend to try even smaller RAM configurations and older processors. A 486 build sits on their workbench. Early results show the same principles apply, though speed falls to fractions of a token per second. The group also explores quantization below 2 bits. Their goal is not speed records but proof that intelligence can live in constrained spaces. The message resonates. Bigger is not always necessary. Smarter compression often suffices.

And so an old beige box from the Clinton era now hums with neural activity. It parses sentences. It offers answers. It does so without fanfare or subsidies from hyperscalers. The demonstration carries quiet defiance. It says the future of AI may not belong solely to those who can afford the latest silicon. With careful engineering, yesterday’s hardware can still contribute. The Pentium II has earned a second act.

Subscribe for Updates

AITrends Newsletter

The AITrends Email Newsletter keeps you informed on the latest developments in artificial intelligence. Perfect for business leaders, tech professionals, and AI enthusiasts looking to stay ahead of the curve.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us