Our paper “Time, Space, and Modality: Probing Earth Foundation Models for Intra-Field Crop Yield Forecasting” (Hamid Kamangir, Heesup Yun, J. Mason Earles) was accepted to the 2nd Workshop on Representation Learning for Earth Observation (REO2) at NeurIPS 2026.

Monthly embeddings projected onto the first two principal components. The mean seasonal path for each country is drawn month by month, from emergence (t1) to harvest (t*).

Earth foundation models are pretrained on large amounts of satellite imagery, and their embeddings are increasingly used as features for crop yield prediction. Most studies report whether these embeddings improve accuracy. We asked two more specific questions:

  1. Do the embeddings follow the crop’s development through the season, so they can be used for in-season forecasting?
  2. Do they keep the spatial detail that separates the best and worst zones within a field, and does that come from multimodal pretraining or from how the model tokenizes space?

Data and models

We compared four Earth foundation models, AlphaEarth (annual, pixel level), OlmoEarth (monthly, patch tokens), Presto (pixel time series), and AnySat (flexible resolution, pretrained on 11 sensors), against a multi-view gated fusion (MVGF) baseline trained from scratch on Sentinel-2, weather, elevation, and soil data. Each time-series model was tested with a simple regression head and with a ConvLSTM head.

The data come from the YieldSAT benchmark: 140 corn field-years (Argentina and Brazil) and 197 wheat field-years (Argentina and Germany). We used 10-fold cross-validation grouped by field so the same field never appears in both training and test sets.

Accuracy over the season

At harvest, OlmoEarth with a ConvLSTM head was the most accurate model for both crops, at both the field and pixel level:

Model Corn field R² Corn pixel R² Wheat field R² Wheat pixel R²
MVGF baseline (no embedding) 0.448 0.244 0.744 0.411
AlphaEarth + regression 0.566 0.397 0.847 0.694
OlmoEarth + ConvLSTM 0.614 0.427 0.863 0.717
Presto + ConvLSTM 0.375 0.276 0.721 0.472
AnySat + ConvLSTM 0.500 0.362 0.811 0.626

Every ConvLSTM head beat the regression head on the same embedding, which means the monthly sequence carries useful information beyond a single summary vector.

Monthly field-level R squared for corn and wheat from seven months before harvest to harvest, for each model
Field-level R² by forecast lead time for (a) corn and (b) wheat. Shaded bands show ±1 standard deviation across the 10 folds; the star marks AlphaEarth's annual embedding.

Over the season, OlmoEarth with ConvLSTM passed AlphaEarth’s annual embedding about four months before corn harvest and about five months before wheat harvest, and stayed ahead until harvest.

The animation at the top helps explain this. OlmoEarth traces a clean seasonal loop. For corn, Argentina and Brazil follow nearly the same loop; for wheat, the loop splits by country, which matches the opposite growing seasons of the two hemispheres. Presto averages its tokens over all months, so its trajectories stay compact and lose the order of the season. AnySat’s trajectories look closer to noise.

Yield maps

The spatial results split between two models. All models smoothed the yield maps compared with the ground truth, but AnySat with ConvLSTM stayed closest on spatial autocorrelation (Moran’s I gap of 0.065 for corn and 0.123 for wheat). OlmoEarth with ConvLSTM matched the yield distribution best (Wasserstein distance of 0.328 for corn and 0.170 for wheat), with AnySat close behind (0.331 and 0.171). The MVGF baseline was worst on both measures.

Ground truth yield map of a corn field in Argentina next to yield maps predicted by OlmoEarth, Presto, and AnySat, all on the same color scale
Ground truth and predicted yield maps for a corn field in Argentina (2021 season, held-out test fold), on a shared color scale. Pixel MAPE, MAE, and R² are shown above each prediction.

The maps above show one held-out corn field. OlmoEarth predicts on a 40 m patch grid, so its map looks blocky, but it places the high- and low-yield zones in the right spots (R² = 0.62, MAE 1.84 t/ha). Presto (R² = 0.11) and AnySat (R² = 0.27) keep the 10 m pixel grid, but their maps are much flatter than the ground truth.

Takeaways

Tracking the season and keeping spatial detail depended on different design choices. OlmoEarth’s monthly attention across space, time, and modality tracked the season and best matched the yield distribution, while AnySat’s broad sensor pretraining with flexible-resolution tokens best kept the spatial autocorrelation. A model for in-field yield forecasting likely needs both: attention over the whole season and fine, flexible spatial tokens.

Next, we plan to extend the evaluation to soybean, rapeseed, and vineyards, and to test fine-tuning and out-of-region splits.