Skip to content
The AI-First Web

NASA and IBM's lunar AI hints at what unlabeled data archives may be worth

One open model, trained on lunar data, beat its rivals on ice prediction. Part of the edge came from design, not data.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
NASA and IBM's lunar AI hints at what unlabeled data archives may be worth

AI-generated image for WebPulse. About our images

In brief
  • NASA and IBM released an open lunar model that beat comparison models on ice prediction, though a control with no lunar pretraining also beat most of them.
  • Large unlabeled archives may reduce the labeling a task needs, but this study cannot say how much of the gain comes from data and how much from design.
  • Ask vendors for results on your own data, with run-to-run variance and known failures reported.

Labels may be the scarcer resource

NASA's new lunar model starts from a problem many organisations will recognise. The team says lunar observation data is plentiful, but labels are scarce. A label is a human-checked answer, such as "this is a crater" or "this region holds ice."

NASA and IBM Research, working with academic partners, have released the NASA-IBM Lunar Foundation Model. The Decoder reports that the two organisations bill it as among the first openly published foundation models for lunar science. A foundation model first learns from huge volumes of unlabeled material. It is then adapted to a specific job using only a few labeled examples.

NASA's chief science data officer, Kevin Murphy, said that "collecting data is only part of the job." Scientists also need the data to be easier to use. Many companies sit on the same kind of asset: years of logs, images or sensor readings that no one has annotated.

17
Years of Lunar Reconnaissance Orbiter observations behind the bulk of the data
Source: NASA and IBM, as reported by The Decoder (October 4, 2026)

How the model was built

The team trained the model from scratch on a corpus called SomBench. It holds nearly 2 million bundles of map tiles, covering 11 data types at two spatial scales.

Two cameras supply the imagery. The Narrow Angle Camera contributes about 1 million sharp pictures, where each pixel spans roughly a meter of ground. The Wide Angle Camera adds just under 964,000 multispectral images, with each pixel spanning 100 meters.

Gravity data from GRAIL, hydrogen data from Lunar Prospector and mineral data from Japan's Kaguya/SELENE probe fill out the set. Altogether the corpus aligns more than 30 layers from nine instruments on four missions.

Two design choices matter for any organisation. First, the team divided training and test data by map zone instead of shuffling tiles at random. That stops near-identical neighbouring tiles from handing the model the answers.

Second, light angle can mask what the ground is actually like. So the team gave the model each tile's sun position, illumination angles and extent. It did not make the model infer them from pixels.

A third choice is architectural. Each data layer gets its own processing path, while baseline models stack all inputs together as channels. The researchers say part of the model's advantage comes from this.

Strong results, and what they do not prove

The biggest gain came in predicting ice in permanently shadowed polar regions. They stay cold enough to preserve water ice and are seen as a possible source of water, oxygen and rocket fuel. On this task, IBM reports error dropping by as much as 22 percent relative to SwinV2-B, the strongest of the comparison models.

22%
Reduction in ice-prediction error versus SwinV2-B, up to
Source: IBM, as reported by The Decoder (October 4, 2026)

Now the catch. The researchers also built a control: the same design, but with randomly set starting weights and no lunar pretraining. On ice prediction, that control still outperformed five of the seven baseline models.

5 of 7
Baselines beaten on ice prediction by the randomly initialized control
Source: technical report, as reported by The Decoder (October 4, 2026)

So some of the edge comes from architecture, not from the pretraining data. The authors also say they have not yet run tests that separate out each design change. A few of the evaluation sets are small.

The pretraining did show a benefit. Per the technical report, the pretrained model matched or beat the random-start control on every task. What remains unclear is how large that gap is.

For crater detection at the 100-meter scale, IBM puts the gain over SwinV2-B at nearly 19 percent, using half the training data. The Decoder reads that as a hint that fewer labeled examples may suffice.

Other results are closer. On meter-scale craters and young volcanic features, the model roughly ties the strongest baselines, within the variance between training runs. IBM cites a 3 percent edge on the volcanic features. The Decoder judges the outcome closer to a draw.

The failures are stated openly. When the model generated maps, its coordinates were sometimes off by several tens of degrees. The authors present it as a base for later tasks, not a replacement for physical instruments.

Nearly 2 million
Tile bundles in the training corpus
Source: NASA and IBM, as reported by The Decoder (October 4, 2026)

What this means for your organisation

This is one research release, not proof of a trend. It does suggest an experiment worth running. Pretraining on unlabeled data may reduce how much labeling a task needs. But this study cannot say how much of the benefit comes from the data and how much from the model's design.

The lunar team also shows habits worth copying. It supplied known context instead of forcing the model to guess it. It kept test data separate by region. And it reported where the model falls short, because errors that large would rule out some uses.

Questions to put to your team: Which of our archives are large, aligned and unlabeled? Which decisions are stalled because labeling is slow or costly? Do we record context, such as time, device or conditions, that a model could use as input? Do related records leak across our test split? And what failure has the vendor not shown us?

Ask for results on your own data, with a control that skips pretraining and the variance between runs reported. A single headline percentage is only a starting point.

The lesson of the lunar model is a modest one. Unlabeled archives may be worth more than they look, but a good control shows how much of the credit belongs to the data.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: The Decoder.

Share this insight