NASA and IBM Release Open-Source Lunar Foundation Model
Trained on nearly 2 million image bundles from 17 years of lunar orbiter data, the open-source vision transformer helps scientists detect polar ice and map rugged terrain.

Key takeaways
- NASA and IBM open-sourced the NASA-IBM Lunar Foundation Model under an Apache-2.0 license on Hugging Face and GitHub.
- The model is pretrained on SomBench, a multimodal dataset of roughly 2 million tile bundles combining 17 years of Lunar Reconnaissance Orbiter data with missions like GRAIL and Kaguya.
- Benchmark results show up to a 22% reduction in error when predicting polar ice stability compared to standard vision baselines.
- Using lightweight LoRA fine-tuning and the TerraTorch toolkit, the architecture lets researchers adapt the system with lower compute requirements.
NASA and IBM Research have released the NASA-IBM Lunar Foundation Model, an open-source artificial intelligence system built to analyze decades of lunar observation data. Announced on September 10, 2026, the model is designed to accelerate planetary science by automating complex mapping tasks, including detecting craters, studying volcanic features, and identifying stable subsurface ice deposits near the Moon's poles.
Pretrained on roughly 2 million image bundles spanning 17 years of orbital observations, the model is available on Hugging Face under an Apache-2.0 license, with supporting code and datasets hosted on GitHub and integrated into the open-source TerraTorch toolkit.

Solving the Data Bottleneck in Lunar Science
For decades, planetary orbiters have accumulated petabytes of surface observations. Data collected by NASA's Lunar Reconnaissance Orbiter (LRO) alone exceeds the volume of all other NASA planetary missions combined. However, analyzing these archives has traditionally required manual evaluation or specialized, single-task algorithms that are computationally intensive to train from scratch.
"NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job," said Kevin Murphy, chief science data officer and acting chief data and AI officer at NASA Headquarters, in the agency announcement. "We also have to make data easier for scientists to explore and use. The NASA-IBM Lunar Foundation Model shows what’s possible when we bring AI to NASA’s petabytes of scientific data."
To build a unified framework, IBM and NASA compiled SomBench, a machine-learning-ready corpus that integrates more than 30 spatially aligned layers across nine instruments from four space missions. The dataset includes roughly 1 million high-resolution panchromatic images at 1 meter per pixel from the LRO Narrow Angle Camera (NAC) and nearly 964,000 multispectral images at 100 meters per pixel from the Wide Angle Camera (WAC). It also incorporates gravity measurements from NASA's GRAIL mission, hydrogen data from Lunar Prospector, and mineralogy maps from JAXA's SELENE/Kaguya orbiter.
Architecture and Training Technical Specifications
According to the model documentation and technical paper, the NASA-IBM Lunar Foundation Model uses a Vision Transformer base (ViT-B) architecture with 12 layers, 12 attention heads, and a 768-dimensional hidden state. It adapts the masked-token training framework from TerraMind, an Earth observation model previously developed by IBM, the European Space Agency (ESA), and Forschungszentrum Jülich.
Training the model required approximately 1,100 GPU-hours across 16 Nvidia H100 GPUs over 150,000 steps using bfloat16 precision. Two distinct adaptations were implemented for lunar conditions:
- Explicit Illumination Geometry: Surface appearance on the airless Moon is heavily distorted by changing sun angles and long shadows. The model ingests optical metadata—such as solar incidence, phase angles, and sub-solar coordinates—as explicit sequence tokens rather than forcing the neural network to infer lighting variations purely from raw pixels.
- Joint Mixed-Resolution Pretraining: Instead of training separate models for different imaging instruments, high-resolution NAC tiles (1 m/pixel) and wider WAC tiles (100 m/pixel) are trained jointly in a single mixed-batch pipeline.
- FlexiViT Patch Resizing: The checkpoint incorporates flexible patch embeddings, allowing fine-tuning at alternate patch resolutions without retraining the entire backbone.

Benchmark Performance Across Core Lunar Tasks
IBM and NASA evaluated the foundation model across several core planetary science tasks, comparing it against general-purpose vision architectures like SwinV2-B, ConvNeXt-B, and ViT-B MAE, according to the model documentation and technical paper.
Polar Ice Prospectivity
Permanently shadowed regions at the lunar poles can remain cold enough to preserve water ice for billions of years, offering potential water, oxygen, and fuel sources for future Artemis missions. In tests predicting polar ice stability from multimodal topography and thermal datasets, the foundation model reduced root-mean-square error (RMSE) by up to 22% compared to an ImageNet-pretrained SwinV2-B baseline (0.0293 vs 0.0377 RMSE).
Crater Mapping and Label Efficiency
Craters provide critical data for dating surface ages and selecting safe landing sites. On context-scale WAC imagery (100 meters per pixel), the foundation model achieved a mean Average Precision (mAP) of 0.2541 when fine-tuned on only 50% of the training data—surpassing the 0.2313 mAP achieved by a SwinV2-B baseline trained on 50% data and the 0.2420 mAP achieved on the full 100% dataset, according to the model card evaluation table. At meter-scale resolution (NAC images), the model performed comparably to top baselines (0.1543 mAP for LoRA adaptation vs 0.1552 for SwinV2-B).
Volcanic Geology
To map irregular mare patches—younger volcanic structures that provide insight into the Moon's thermal history—the model reached an intersection over union ($IoU_1$) of 0.5709 with a frozen encoder, outperforming baseline models and demonstrating higher parameter efficiency.
Fine-Tuning Efficiency and Operational Scope
The researchers highlighted that lightweight Low-Rank Adaptation (LoRA) on encoder attention and MLP layers matched or exceeded full fine-tuning performance across most benchmarks while training only a small fraction of the total parameters.
However, the model card specifies clear boundaries for operational use. The model does not maintain an internal geodetic reference frame, meaning generated outputs can experience absolute coordinate or elevation drifts and are not certified for real-time hazard clearance or automated landing-site selection. Instead, the model is intended to function as an open scientific framework for dense regression, semantic segmentation, and planetary feature discovery.
Frequently asked questions
Where is the NASA-IBM Lunar Foundation Model available?
The model checkpoints and tokenizers are hosted on Hugging Face under the Apache-2.0 license, while training scripts, fine-tuning tools, and benchmark datasets are available via GitHub and the TerraTorch library.
What datasets were used to train the lunar model?
The model was trained from scratch on SomBench, an aggregation of roughly 2 million tile bundles combining 17 years of NASA Lunar Reconnaissance Orbiter imagery with data from NASA's GRAIL, Lunar Prospector, and JAXA's SELENE/Kaguya missions.
How does the model handle extreme lunar shadows and lighting?
Rather than forcing the model to infer lighting solely from image pixels, the architecture tokenizes explicit optical metadata—such as solar incidence angles, emission, and tile footprints—directly into the encoder.
Can the model be used directly for autonomous spacecraft landings?
No. The model card states that the system does not maintain an absolute geodetic reference frame and is designed for scientific analysis and mapping rather than operational landing-site hazard certification.
Sources
- NASA, IBM Launch AI Foundation Model for Lunar Science - NASA ScienceNASA Science · Sep 10, 2026 · Official
- IBM and NASA Release Open-Source AI Model to Support Lunar ExplorationIBM Newsroom · Official
- A rough guide for going back to the MoonIBM · Sep 10, 2026 · Official
- nasa-ibm-ai4science/NASA-IBM-Lunar-Foundation-ModelHugging Face · Sep 1, 2026 · Official
- NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar scienceThe Decoder · Oct 4, 2026
- NASA, IBM launch new AI model for studying the moonSpace · Sep 12, 2026
- IBM and NASA release an open-source lunar foundation modelTNW | Space · Sep 10, 2026
How this story was made: the newsroom picked it up from the-decoder.com, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (24 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 5, 2026 at 00:38 UTC


