Crop Phenology in Monthly Earth Foundation Model Embeddings
Our paper “Time, Space, and Modality: Probing Earth Foundation Models for Intra-Field Crop Yield Forecasting” (Hamid Kamangir, Heesup Yun, J. Mason Earles) was accepted to the 2nd Workshop on Representation Learning for Earth Observation (REO2) at NeurIPS 2026.
Earth foundation models are pretrained on large amounts of satellite imagery, and their embeddings are increasingly used as features for crop yield prediction. Most studies report whether these embeddings improve accuracy. We asked two more specific questions:
- Do the embeddings follow the crop’s development through the season, so they can be used for in-season forecasting?
- Do they keep the spatial detail that separates the best and worst zones within a field, and does that come from multimodal pretraining or from how the model tokenizes space?
Data and models
We compared four Earth foundation models, AlphaEarth (annual, pixel level), OlmoEarth (monthly, patch tokens), Presto (pixel time series), and AnySat (flexible resolution, pretrained on 11 sensors), against a multi-view gated fusion (MVGF) baseline trained from scratch on Sentinel-2, weather, elevation, and soil data. Each time-series model was tested with a simple regression head and with a ConvLSTM head.
The data come from the YieldSAT benchmark: 140 corn field-years (Argentina and Brazil) and 197 wheat field-years (Argentina and Germany). We used 10-fold cross-validation grouped by field so the same field never appears in both training and test sets.
Accuracy over the season
At harvest, OlmoEarth with a ConvLSTM head was the most accurate model for both crops, at both the field and pixel level:
| Model | Corn field R² | Corn pixel R² | Wheat field R² | Wheat pixel R² |
|---|---|---|---|---|
| MVGF baseline (no embedding) | 0.448 | 0.244 | 0.744 | 0.411 |
| AlphaEarth + regression | 0.566 | 0.397 | 0.847 | 0.694 |
| OlmoEarth + ConvLSTM | 0.614 | 0.427 | 0.863 | 0.717 |
| Presto + ConvLSTM | 0.375 | 0.276 | 0.721 | 0.472 |
| AnySat + ConvLSTM | 0.500 | 0.362 | 0.811 | 0.626 |
Every ConvLSTM head beat the regression head on the same embedding, which means the monthly sequence carries useful information beyond a single summary vector.
Over the season, OlmoEarth with ConvLSTM passed AlphaEarth’s annual embedding about four months before corn harvest and about five months before wheat harvest, and stayed ahead until harvest.
The animation at the top helps explain this. OlmoEarth traces a clean seasonal loop. For corn, Argentina and Brazil follow nearly the same loop; for wheat, the loop splits by country, which matches the opposite growing seasons of the two hemispheres. Presto averages its tokens over all months, so its trajectories stay compact and lose the order of the season. AnySat’s trajectories look closer to noise.
Yield maps
The spatial results split between two models. All models smoothed the yield maps compared with the ground truth, but AnySat with ConvLSTM stayed closest on spatial autocorrelation (Moran’s I gap of 0.065 for corn and 0.123 for wheat). OlmoEarth with ConvLSTM matched the yield distribution best (Wasserstein distance of 0.328 for corn and 0.170 for wheat), with AnySat close behind (0.331 and 0.171). The MVGF baseline was worst on both measures.
The maps above show one held-out corn field. OlmoEarth predicts on a 40 m patch grid, so its map looks blocky, but it places the high- and low-yield zones in the right spots (R² = 0.62, MAE 1.84 t/ha). Presto (R² = 0.11) and AnySat (R² = 0.27) keep the 10 m pixel grid, but their maps are much flatter than the ground truth.
Takeaways
Tracking the season and keeping spatial detail depended on different design choices. OlmoEarth’s monthly attention across space, time, and modality tracked the season and best matched the yield distribution, while AnySat’s broad sensor pretraining with flexible-resolution tokens best kept the spatial autocorrelation. A model for in-field yield forecasting likely needs both: attention over the whole season and fine, flexible spatial tokens.
Next, we plan to extend the evaluation to soybean, rapeseed, and vineyards, and to test fine-tuning and out-of-region splits.