NASA and IBM Research have released the NASA-IBM Lunar Foundation Model, an open-source tool designed to make decades of lunar observation data usable for machine learning. The model, developed in collaboration with several academic institutions, is described by the organizations as one of the first open-source foundation models specifically for lunar science.

What Happened

Kevin Murphy, NASA's chief science data officer, noted that "NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job." He added that the data also has to be easier for scientists to use. To address this, the team trained the model from scratch using SomBench, which the team states is the largest co-registered multimodal lunar corpus to date. The dataset comprises nearly 2 million tile bundles across 11 modalities and two spatial scales, drawing primarily from 17 years of observations by the Lunar Reconnaissance Orbiter (LRO). It also incorporates data from the GRAIL mission, Lunar Prospector, and JAXA's Kaguya/SELENE probe, combining over 30 spatially aligned data layers from nine instruments.

Unlike task-specific algorithms, this foundation model is pretrained on large volumes of unlabeled data and can be adapted to specific tasks with limited labeled examples. The architecture is based on TerraMind, a multimodal Earth observation model, but the researchers trained the lunar version from scratch rather than fine-tuning it. A key technical feature is the explicit input of imaging geometry, including illumination angles and sun position, allowing the model to account for lighting conditions that significantly affect lunar surface appearance.

Why It Matters

The release addresses a critical bottleneck in lunar research: the abundance of observation data contrasted with the scarcity of labeled datasets. By providing a reusable foundation, the model enables scientists to adapt it to various tasks with fewer labeled examples. IBM reports that the model demonstrated significant improvements in predicting polar ice deposits, cutting prediction error by up to 22 percent compared to the best baseline, SwinV2-B. Permanently shadowed polar regions are considered potential resources for water, oxygen, and rocket fuel.

Part of the model's performance edge stems from its architecture, which assigns each data layer its own processing path. This approach contrasts with baseline models that treat all inputs as stacked channels, allowing the system to better handle the distinct characteristics of different sensor data. The model also showed a nearly 19 percent improvement over baselines in coarse-scale crater detection when trained with only half the data, suggesting high data efficiency. However, the researchers note limitations, including unsuitability for absolute geodetic positioning and the fact that controlled experiments isolating specific innovations are still pending. The model is part of the broader NASA-IBM "AI for Science" collaboration, which has previously released the Prithvi model for Earth observation tasks.

The Bottom Line

The NASA-IBM Lunar Foundation Model is publicly available on Hugging Face, with code hosted on GitHub and integrated into the TerraTorch toolkit. While it offers a powerful new tool for analyzing lunar data—particularly for ice detection and crater mapping—the authors emphasize it is a foundation for downstream tasks rather than a replacement for physical measurement instruments.