Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis

2026-08-20 · Hugging Face

Introducing OlmoEarth Embeddings: Custom Embedding Exports from OlmoEarth Studio for Downstream Analysis

OlmoEarth Studio, AllenAI's platform for building Earth observation models, now supports computing and exporting embedding vectors. These are compact numerical representations of Earth-observation data produced by the open-source OlmoEarth foundation models. The source code and model weights are publicly released alongside the research paper, enabling the community to fully inspect how the embeddings are generated.

Embeddings offer a fast, cost-effective entry point for leveraging OlmoEarth. They support a broad range of downstream tasks including similarity search, segmentation, and unsupervised exploration. Locations sharing similar surface characteristics map to similar vectors, while dissimilar locations map far apart in embedding space. OlmoEarth embeddings have demonstrated strong performance both in AllenAI's internal benchmarking and in independent evaluations.

The exported files are lightweight Cloud-Optimized GeoTIFFs (COGs) that are easy to share and use. Users select their area of interest, time range, encoder variant, spatial resolution, and imagery sources via the Studio UI or API, and receive a COG ready for any downstream application. For applications requiring even higher performance, Studio also supports supervised fine-tuning (SFT).

Custom-computed embeddings are now available to OlmoEarth Studio users. Interested parties should reach out for access. Instructions for computing embeddings locally with the publicly released OlmoEarth models are also provided.

Computing Embeddings in Studio

The workflow for computing embeddings is identical to any other prediction task in Studio: configure a model, execute the job, then download the output. Several parameters allow precise tailoring of results:

  • Area of interest: Draw a polygon or upload one; Studio automatically acquires and tiles the necessary imagery.
  • Time span: Select between 1 and 12 monthly periods.
  • Encoder variant: Nano (128-dimensional, 1.4M parameters), Tiny (192-dimensional, 6.2M parameters), or Base (768-dimensional, 89M parameters).
  • Spatial resolution: 10m, 20m, 40m, or 80m per pixel.
  • Imagery sources: Sentinel-2 L2A, Sentinel-1 RTC, or both combined.

Studio returns a COG containing one band per embedding dimension. The vectors are quantized and stored as signed 8-bit integers (int8), ranging from -127 to +127, with -128 reserved for no-data pixels. Floating-point vectors can be recovered using the `dequantize_embeddings` function provided in `olmoearth_pretrain`.

Because embeddings are generated on-demand rather than retrieved from a pre-computed global archive, they exactly match the user's specified area, time window, and sensor conditions. This enables creation of monthly embeddings that capture seasonal dynamics instead of being limited to annual composites.

What You Can Do with OlmoEarth Embeddings

All examples below use the OlmoEarth-v1-Tiny (192-dim) encoder at 40-meter resolution with Sentinel-2 L2A composites (annual for most cases, monthly for change detection). Although Tiny is a lightweight model, it remains highly performant. Users may substitute larger variants when greater capacity is required, at the expense of increased compute and storage.

Similarity Search: Finding "More Like This"

Select any query pixel (or the mean embedding of a small window), extract its vector, and compute cosine similarity against every other pixel in the scene. The resulting heatmap immediately reveals which locations in the landscape are most and least similar to the query.

When the query is taken near the urban center of Merced, California, urban fabric and road corridors light up coherently while agricultural parcels remain dark. The model distinguishes built-up surfaces from cropland without any supervision or labels.

Switching the query to a small agricultural window and using the mean vector of that window, the highest-similarity patches (cosine similarity ≥ 0.89) are all irrigated agricultural fields. The lowest-similarity locations (similarity near zero) include an airport with surrounding bare ground, a reservoir with dry terrain, and arid rangeland. No training data or labels are used — only a dot product in embedding space.

Few-shot Segmentation: Labeling the Landscape

While similarity search answers "where is it like this?", discrete wall-to-wall maps sometimes require labeled categories. Because OlmoEarth embeddings already contain rich semantic information, a simple linear classifier trained on very few labeled pixels can produce accurate land-cover maps across an entire region.

In a test over Ca Mau, Vietnam — a coastal mangrove area — only 60 pixels were labeled (20 pixels per class) using ESA WorldCover 2021 as ground truth for three classes: mangrove, water, and other. After randomly sampling these pixels, a logistic regression classifier with per-feature standardization was trained and used to predict every pixel in the scene.

From just 60 labeled pixels, the classifier produced a coherent map with a weighted F1 score of 0.84. Mangrove stands, tidal channels, and open water bodies were clearly delineated across the whole region. Performance saturates quickly: increasing the label count from 30 to 300 produced almost no accuracy gain, indicating that the embeddings themselves perform most of the heavy lifting.

The core analysis requires only a few lines of Python code:


import rasterio
import numpy as np
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

# Load the 192-band embedding COG exported from Studio
with rasterio.open("embeddings.tif") as ds:
    emb = ds.read().astype(np.float32) # (192, H, W)
C, H, W = emb.shape
X = emb.reshape(C, -1).T # (H*W, 192)

# Train on labeled pixels, predict everywhere
clf = make_pipeline(StandardScaler()

Source