AgriEmbedding
Technical Evaluation Report
2026‑08‑13
Technical Evaluation Report

Global and Dense Visual Representations

Retrieval quality, patch-level representation quality, intended uses, and evaluation limitations for the two representation outputs of AgriEmbedding (a.k.a SEED-EMBEDDINGS) .

Model AgriEmbedding CLS-token global retrieval Patch-token dense representation

Abstract

AgriEmbedding is a vision embedding model that exposes two representation interfaces from the same forward pass: a CLS-token output, a single pooled vector per image, used for global, whole-image retrieval, and a patch-token output, a grid of per-patch vectors, used for dense, local representation tasks.

We evaluate each output on the benchmark that suits it: the CLS token on Image→Image retrieval (Public and AgriStress-500 V1.0.2 suites), and the patch tokens on feature separation on AgriStress-500 V1.0.2, at coarse (crop/weed/soil) and fine-grained (individual crop and weed subtypes) label granularity.

On retrieval, AgriEmbedding leads every external baseline we scored, on both suites; Section 5 lists them in full. The closest competitor differs by suite: PlantCLEF2024 on Public, 0.18 pp behind, and BioCLIP on AgriStress-500 V1.0.2, 0.90 pp behind. We compute no confidence intervals for retrieval, so read the Public margin as a tie (Section 5).

The dense benchmark scores four of these baselines, TIPSv2-base, TIPSv2-LARGE, DINOv2-LARGE and PlantCLEF2024, from the same checkpoints, and the result depends on label granularity. At crop-subtype granularity AgriEmbedding leads all four on k-NN mIoU, k-NN accuracy, ARI and NMI. At weed-subtype granularity it leads on k-NN accuracy, ARI, NMI and both labels that separate, while TIPSv2-base takes macro k-NN mIoU by 0.13 pp. At coarse granularity all four lead it on k-NN mIoU and linear-probe mIoU, and it is third of the five on k-NN accuracy and on adjacent-patch smoothness.

Summary of results

What subtype-level weed separation enables

Weeds compete with the crop for water, nutrients, and light. The conventional response is to broadcast herbicide across the whole field, which contributes to herbicide resistance. Targeted application instead requires species-level identification, so the herbicide reaches the weed present rather than the whole field. Resistance management depends on it too, since the species that drive multi-herbicide resistance in row crops have to be identified before they can be treated differently from the rest of the field. Coarse crop/weed/soil masks do not support any of this; subtype separation does.

Whole-image similarity
CLS tokenImage→Image retrieval, nDCG@10
Separability, fine-grained
Patch tokensk-NN & linear-probe mIoU per subtype
Separability, coarse
Patch tokensk-NN & linear-probe mIoU
Unsupervised structure
Patch tokenssilhouette in Section 6.5; ARI and NMI in Appendix B
Spatial coherence
Patch tokensAdjacent-patch similarity

Principal findings

  • Retrieval: AgriEmbedding leads every external baseline we scored on both suites: 82.32% on Public, 73.95% on AgriStress-500 V1.0.2. The closest competitor differs by suite: PlantCLEF2024 on Public (82.14%, 0.18 pp behind), BioCLIP on AgriStress-500 V1.0.2 (73.05%, 0.90 pp behind).

  • Fine-grained crop subtypes: AgriEmbedding leads all four external baselines on k-NN mIoU (76.64%), k-NN accuracy, ARI, and NMI, with TIPSv2-base closest at 75.17%. TIPSv2-base leads linear-probe mIoU; PlantCLEF2024 is fourth of the five on k-NN mIoU, k-NN accuracy, and NMI.

  • Fine-grained weed subtypes: two of the four labels separate at all, grass (0.570 k-NN mIoU) and the generic weed class (0.904). AgriEmbedding is highest of the five models on both, mean 0.737 across the two, and it also leads k-NN accuracy (90.67%), ARI (0.081) and NMI (0.105) over the full label set. TIPSv2-base takes macro k-NN mIoU, 21.43% against 21.30%, a margin far inside both confidence intervals.

  • Spatial coherence: AgriEmbedding's adjacent-patch similarity (0.866) exceeds DINOv2-LARGE (0.671) and PlantCLEF2024 (0.581, lowest of the five), and is third of the five behind TIPSv2-LARGE (0.902) and TIPSv2-base (0.871).

  • Dense, coarse: AgriEmbedding scores 57.91% k-NN mIoU, 56.13% linear-probe mIoU and 80.90% k-NN accuracy. All four baselines beat it on k-NN mIoU: TIPSv2-base (61.16%), TIPSv2-LARGE (60.69%), PlantCLEF2024 (59.74%) and DINOv2-LARGE (59.21%); TIPSv2-LARGE has the highest linear-probe mIoU of the five (61.42%).

Decision guidance

Which output to use, by use case.

Whole-image retrieval / similar-image search
CLS-token
Leads every baseline we scored on both suites; margins and the Public tie in Section 5, retrieval
Frozen-feature coarse dense classification
Patch-token
Only representation evaluated for this task; see Section 6.4, coarse granularity, for standing against TIPSv2-base, TIPSv2-LARGE, DINOv2-LARGE and PlantCLEF2024
Frozen-feature fine-grained crop-subtype classification
Patch-token
Leads all four baselines on k-NN mIoU, k-NN accuracy, ARI and NMI (Section 6.1)
Spatial heatmaps / dense correspondence
Patch-token
With caveats: third-highest smoothness of the five models, no boundary-level evidence yet
01

Introduction

This is an evaluation report. It documents which KPIs we use, why we chose them, how we measure them, what the results mean, how the CLS-token and patch-token outputs differ, and which conclusions the evidence supports, leaves unsupported, or leaves uncertain. It does not cover how we built the model. Throughout, we means the Precision AI team that ran these evaluations; Precision AI also develops AgriEmbedding and publishes the AgriStress-500 V1.0.2 benchmark used here.

02

Scope and disclosure boundary

We score both representation outputs here: the CLS-token output on Image→Image retrieval, and the patch-token output (the grid described in Section 3) on coarse and fine-grained feature separation. We do not score either on the other's benchmark, since a single pooled vector cannot support per-patch dense evaluation and a per-patch grid cannot rank whole-image similarity.

AgriStress-500 V1.0.2 backs the entire dense benchmark (Section 6) and one of the two retrieval suites (Section 5). We publish it ourselves, under CC-BY-NC-4.0, so anyone can inspect the benchmark and its labels independently. Appendix C documents its composition.

Disclosure boundary

This report covers the representations AgriEmbedding exposes and how they behave on our benchmarks. It does not disclose the training datasets, training objectives, optimization procedure, model lineage, or training compute, and no such details should be inferred from it. Vision-encoder parameter counts (Sections 5 and 6) and patch size (Appendix D) are disclosed where the size-matching discussion requires them; from these two figures a reader could infer an approximate architecture family, which we neither confirm nor elaborate on. AgriStress-500 was never used in any supervised or unsupervised training, retraining, or fine-tuning of any model described in this report, and was assembled and selected independently of model development to reflect field conditions.

03

Representation interfaces

One forward pass, two outputs: a pooled CLS-token vector and a grid of patch-token vectors, each suited to a different downstream task.

side-by-side comparison
PropertyCLS tokenPatch tokens
Output shape1 vectorN vectors (one per patch)
Information retainedOverall content of the imageLocal detail and spatial position
Main use casesWhole-image retrieval, dedup, coarse clusteringDense correspondence, localization, segmentation input
ComparisonCosine similarity against other whole-image vectorsPatch-to-patch, or pooled per region
Computational costOne dot product / comparisonUp to N×N dot products / comparison
Main limitationDiscards within-image spatial detail; cannot localize a matchCost scales with patch count; boundary accuracy not evaluated in this revision
Primary benchmarkImage→Image retrieval, Section 5Subtype & coarse feature separation, Section 6
04

Evaluation framework

We score both tracks on frozen representations with no task head. Retrieval runs under a fixed similarity protocol; dense evaluation runs under nonparametric, linear, and unsupervised probes.

4.1 · primary and supporting KPIs
CapabilityPrimary KPISupporting KPIs
RetrievalnDCG@10Precision@1/@5/@10, nDCG@1/@5, MAP@10
Dense feature utilityk-NN mIoUk-NN accuracy, linear-probe mIoU
Unsupervised structureNone (exploratory)Class-label silhouette in Section 6.5; ARI and NMI tabulated in Appendix B
Spatial behaviorNone (exploratory)Adjacent-patch similarity

We treat no single KPI as a complete measure of model quality; each one is scoped to the question it answers.

4.2 · retrieval KPI rationale
KPIProduct question
Precision@kIf a user only looks at the first k results, what fraction are relevant?
nDCG@kHow close to rank 1 are the relevant results?
MAP@10Does the model surface every relevant item early?
4.3 · dense KPI rationale
KPIProduct question
k-NN mIoUDoes raw feature geometry separate classes with no classifier?
Linear-probe mIoUIs the feature space linearly separable by class?
ARI / NMIDoes class structure emerge with zero label access?
Class-label silhouetteHow cleanly do true classes separate geometrically?
Adjacent-patch similarityAre neighboring patches spatially consistent?
05

CLS-token retrieval benchmark

We test whether AgriEmbedding's CLS-token representation ranks truly related images above unrelated ones, given a query image.

We compare AgriEmbedding (22M, via its CLS-token output) against every external baseline that exposes a pooled image embedding, listed below in full with vision-encoder parameter counts. We score Image→Image retrieval on two suites: Public (built from the CropAndWeed dataset[1], spanning 8 crop types; see Appendix C) and AgriStress-500 V1.0.2 (published by Precision AI; 451 drone images from working fields, spanning 10 crop types and 5 weed-related labels).

models compared, retrieval track
ModelFamily / notesParams
AgriEmbeddingThe model under evaluation22M
TIPSv2-LARGE (baseline)Same checkpoint as Section 6's dense benchmark~303M
TIPSv2-base (baseline)Same checkpoint as Section 6's dense benchmark86M
TIPSv2-SO (baseline)SoViT-400m/14 TIPSv2 configuration412M
TIPSv2-Giant (baseline)ViT-g/14 TIPSv2 configuration1.1B
DINOv2-LARGE (baseline)Same checkpoint as Section 6's dense benchmark300M
PlantCLEF2024 (baseline)Same checkpoint as Section 6's dense benchmark; plant-identification-oriented86M
BioCLIP (baseline)ViT-B/16 biology-domain CLIP model86M
AgriCLIP (baseline)DINO ResNet-50 image encoder + CLIP ViT-B/16 text encoder25.6M
MaskDINO Swin-L (PAI-Finetuned)Detection-transformer backbone, further trained on similar domain data as the AgriEmbedding model197M
Qwen3-VL-Embedding-2B (baseline)General-purpose vision-language embedding model2B
Qwen3-VL-Embedding-8B (baseline)General-purpose vision-language embedding model8B
CropVLM (baseline)ViT-Base image encoder86M
CLIP-L/14 (baseline)General-purpose CLIP vision encoder, ViT-L/14~304M
SigLIP-L/16 (baseline)General-purpose SigLIP image tower, ViT-L/16~303M
EUPE-ViT-T (baseline)ViT-Tiny image encoder, the smallest model in this benchmark6M
Baselines and checkpoints, retrieval track

We score TIPSv2-LARGE, TIPSv2-base, DINOv2-LARGE and PlantCLEF2024 from the same checkpoints we use in the dense benchmark (Section 6); every other baseline above is scored only here, for the reasons in Section 6. Every figure above comes from a public model card or published architecture: the two Qwen sizes from their names, BioCLIP as a published ViT-B/16, TIPSv2-base and TIPSv2-SO from their Hugging Face model cards, AgriCLIP and MaskDINO from the standard counts of their component architectures (a DINO ResNet-50 image encoder and a Swin-L backbone), TIPSv2-Giant as a published 1.1B vision tower, and CLIP-L/14 as a standard ViT-L/14 vision encoder, none of them from figures disclosed for those specific checkpoints. The last three follow the same rule: SigLIP-L/16 takes the 303M published for the SigLIP Large image tower, CropVLM reads as a ViT-Base image encoder, and EUPE-ViT-T reads as a ViT-Tiny. Every row now carries a count. Two of them sit close to our own size: AgriCLIP at ~25.6M is within about 16% of AgriEmbedding, and EUPE-ViT-T is roughly 3.6x smaller and is the only baseline in this report smaller than the model under evaluation. Every other retrieval baseline is at least 3.9x larger.

Each suite has its own roster, and reconciling them is open work

We score each suite against the models we hold results for on that suite, set independently of the other, and we make no attempt to force one roster onto both. Sixteen models have an AgriStress-500 V1.0.2 score; twelve have a Public one. Bringing both rosters to the same models is open work. We welcome anyone who wants to run either benchmark alongside us: AgriStress-500 V1.0.2 is public, and the full protocol is in Appendix D.

Image→Image retrieval on Public: nDCG@10
Image-to-image, ordered highest to lowest. Eleven models carry four significant figures from our own full-precision logs; TIPSv2-Giant is read from the published results at three decimals and is labelled at that precision.
Image→Image retrieval on AgriStress-500 V1.0.2: nDCG@10
Image-to-image, ordered highest to lowest. Eleven models carry four significant figures from our own full-precision logs; TIPSv2-Giant, CropVLM, CLIP-L/14, SigLIP-L/16 and EUPE-ViT-T are read from the published results at three decimals and are labelled at that precision.
Interpretation

AgriEmbedding leads every external baseline we scored on both suites, but the closest competitor differs by suite. On Public, PlantCLEF2024 is closest (82.14% vs. 82.32%, 0.18 pp behind), close enough that the two should be treated as effectively tied. On AgriStress-500 V1.0.2, BioCLIP is closest (73.05% vs. 73.95%, 0.90 pp behind), with PlantCLEF2024 third (70.43%). Below the near-tie the field opens up. On Public the next model is TIPSv2-Giant at 76.1%, 6.22 pp behind, and the panel runs down from there; on AgriStress-500 V1.0.2 it runs from PlantCLEF2024 at 70.43% downward. We compute no confidence interval for any retrieval metric, so both closest-competitor margins are point estimates only, and they are the tightest margins anywhere in this report.

06

Patch-token dense benchmark

We test whether the raw patch-token feature map, not the pooled CLS output, separates individual crop and weed subtypes and stays spatially coherent.

We compare five models throughout: AgriEmbedding (via its patch-token output), and four external baselines: TIPSv2-LARGE, TIPSv2-base, DINOv2-LARGE, and PlantCLEF2024. We score all five under one identical protocol (Appendix D), against AgriStress-500 V1.0.2 field imagery. Fine-grained subtype results come first (6.1–6.2), then spatial behavior (6.3), then the coarse split (6.4–6.5).

Baselines and checkpoints, dense track

TIPSv2-base, TIPSv2-LARGE, DINOv2-LARGE and PlantCLEF2024 are the four models scored in both benchmarks, from the same checkpoints in each, so a reader can carry them between Section 5 and Section 6. TIPSv2 appears at two sizes; we hold no patch-token log for a smaller configuration of DINOv2-LARGE or PlantCLEF2024. All four expose a patch-token grid at the same patch size as AgriEmbedding, and we run all four frozen, with no task head, through the protocol in Appendix D. PlantCLEF2024 is a DINOv2 ViT-B/14 fine-tuned on plant imagery; we use it as a frozen feature extractor and discard its classification head, so its taxonomic label space plays no part in these results. The comparison is not size-matched: every baseline here is at least 3.9x larger than AgriEmbedding at 22M, and TIPSv2-LARGE and DINOv2-LARGE are roughly 14x larger.

Which models are eligible for this benchmark

This benchmark uses a separate roster from Section 5: a model must expose a patch-token feature map this protocol can score, and we must hold a patch-token log for it. The five models above clear both conditions. Section 5 is broader because a pooled image embedding is a substantially weaker requirement, and clearing that bar does not qualify a model for this one.

AgriCLIP, MaskDINO and both Qwen3-VL-Embedding sizes are ruled out architecturally: AgriCLIP's image encoder is a DINO ResNet-50 with no patch grid to sample, MaskDINO's Swin-L backbone is windowed and hierarchical and is pooled to a single vector before its output layer, and the Qwen models are built on a comparable non-interpolable backbone family. TIPSv2-SO, TIPSv2-Giant, BioCLIP, CLIP-L/14, SigLIP-L/16, CropVLM and EUPE-ViT-T carry no known architectural limit, since all seven are plain patch-grid vision transformers, but we hold no patch-token log for any of them in the source data for this revision. They stay open for future evaluation.

6.1 · Fine-grained separation: crop subtypes

10 individual crop subtypes (barley, canola, chickpea, corn, flax, lentil, oat, peas, soybean, wheat), all with enough samples to be scored, rather than the single "crop" bucket.

k-NN mIoU, k-NN accuracy, linear-probe mIoU: fine-grained crop subtypes
per-crop-type k-NN mIoU (all 10 subtypes eligible)
ModelBarleyCanolaChickpeaCornFlaxLentilOatPeasSoybeanWheat
AgriEmbedding0.6810.8610.6170.7970.7760.9100.6720.8530.7740.723
TIPSv2-base0.6130.8580.6820.7090.8220.9350.5860.8320.7970.683
TIPSv2-LARGE0.5360.8200.5990.6840.7690.8670.5400.7670.7500.629
DINOv2-LARGE0.4110.7280.5000.6430.6130.7090.3410.5610.5560.515
PlantCLEF20240.4630.7440.6420.7500.6930.8020.4180.6710.6790.582
highest of the five models for that column
AgriEmbedding leads on 4 of 5 metrics

Highest k-NN mIoU (76.64% against TIPSv2-base's 75.17%), highest k-NN accuracy (86.68%), highest ARI (0.386, about two and a half times the runner-up), and highest NMI (0.557, about double the runner-up) of the five models. TIPSv2-base leads only linear-probe mIoU (83.51%), a tenth of a point ahead of TIPSv2-LARGE (83.43%). Per crop type, AgriEmbedding is highest on 6 of 10 subtypes and TIPSv2-base on the other 4 (chickpea, flax, lentil, soybean); none of TIPSv2-LARGE, DINOv2-LARGE and PlantCLEF2024 leads any individual crop type. The k-NN mIoU margin over TIPSv2-base is 1.5 pp and the two confidence intervals overlap heavily (Appendix B), so read this as a consistent ordering across metrics rather than a separated one. This is discrimination within a coarse class, telling barley from oat rather than plant from soil. Read absolute levels here as an upper bound, not a generalization estimate.

6.2 · Fine-grained separation: weed subtypes

Weed labels clear the benchmark's sample threshold: three named species plus a generic "weed" catch-all. Two of them separate at all, so the per-label table reports those two. Broadleaf and palmer amaranth score at or near zero IoU for every model, so we do not report them per label.

k-NN mIoU, k-NN accuracy, linear-probe mIoU: fine-grained weed subtypes
Macro-averaged over all four eligible labels, as the benchmark reports them, including the two that separate for no model. Per-label figures are in the table below.
per-species k-NN mIoU, and the mean across the two reported labels
ModelGrassWeed (generic)Mean, 2 labels
AgriEmbedding0.5700.9040.737
TIPSv2-base0.5470.8930.720
TIPSv2-LARGE0.4720.8810.677
DINOv2-LARGE0.3630.8680.615
PlantCLEF20240.4860.8830.684
AgriEmbedding leads on both labels that separate

AgriEmbedding is highest of the five models on both labels that separate, at 0.570 and 0.904 k-NN mIoU, a mean of 0.737 across the two against 0.720 for TIPSv2-base, 0.684 for PlantCLEF2024, 0.677 for TIPSv2-LARGE and 0.615 for DINOv2-LARGE. It also leads k-NN accuracy (90.67%), ARI (0.081) and NMI (0.105). TIPSv2-base takes macro k-NN mIoU, 21.43% against 21.30%; the two confidence intervals overlap almost entirely, so we read that pair as tied. Those macro figures are the benchmark's own; we derive the mean above ourselves, from the two per-label IoUs.

6.3 · Spatial behavior

Smoothness: AgriEmbedding
0.866
Smoothness: TIPSv2-base
0.871
Smoothness: TIPSv2-LARGE
0.902
Smoothness: DINOv2-LARGE
0.671
Smoothness: PlantCLEF2024
0.581

Adjacent-patch similarity is a property of the patch grid, so it does not change with label granularity. AgriEmbedding shows far greater local feature continuity than DINOv2-LARGE (0.866 vs. 0.671) and PlantCLEF2024 (0.581, lowest of the five), and is third of the five behind TIPSv2-LARGE (0.902) and TIPSv2-base (0.871). Smoothness alone does not establish semantic boundary accuracy; separating coherence from oversmoothing requires boundary-aware evaluation, not included in this revision.

Coarse granularity is not what this model targets

The two subsections below score the 3-way crop/weed/soil split, reported for completeness rather than as a target metric. AgriEmbedding is specialized to widen the margin between labels inside a well-represented cluster (6.1 and 6.2), and a 3-way soil/crop/weed mask is a coarser question than it is built to answer. A general-purpose backbone does comparatively better here.

6.4 · Feature quality against labels (coarse)

k-NN mIoU, k-NN accuracy, linear-probe mIoU: coarse

At the 3-way split, TIPSv2-base leads k-NN mIoU and k-NN accuracy, and TIPSv2-LARGE leads linear-probe mIoU. AgriEmbedding is lowest of the five models on both k-NN mIoU and linear-probe mIoU, 3.3 pp behind TIPSv2-base and 5.3 pp behind TIPSv2-LARGE respectively, with TIPSv2-LARGE second on k-NN mIoU (60.69%) and DINOv2-LARGE second on linear-probe mIoU (60.70%). On k-NN accuracy AgriEmbedding is third of the five (80.90% against 81.20%, 80.99%, 80.28% and 79.99%), where the five models sit within 1.2 pp of each other. The ranking is the reverse of 6.1.

6.5 · How cleanly classes separate (coarse)

Silhouette: AgriEmbedding
0.088
Silhouette: TIPSv2-base
0.031
Silhouette: TIPSv2-LARGE
0.023
Silhouette: DINOv2-LARGE
0.025
Silhouette: PlantCLEF2024
0.053

Class-label silhouette, computed against ground-truth class labels rather than cluster assignments, scaled to the highest of the five values. AgriEmbedding is highest, PlantCLEF2024 second. We tabulate coarse ARI and NMI for all five models in Appendix B rather than charting them here.

Supported conclusion

The strongest dense evidence in this report is fine-grained subtype separation from frozen patch tokens. AgriEmbedding leads all four baselines on k-NN mIoU, k-NN accuracy, ARI and NMI at crop-subtype granularity, and on both weed labels that separate at all, though its margin over TIPSv2-base sits inside the confidence intervals. It is third of the five on spatial coherence, behind TIPSv2-LARGE and TIPSv2-base. The evidence does not support a general claim about unsupervised segmentation, or production readiness for the sample-starved weed species. At the coarse split (6.4–6.5) all four baselines lead on the metrics that use labels.

07

Claim-to-evidence matrix

Every claim this report makes, with the evidence behind it. We evaluate retrieval and dense capability on different benchmarks against their own baselines, and we do not compare them against each other. We report coarse-granularity standing in Sections 6.4 and 6.5.

ClaimEvidenceScopeConfidence
Retrieval capability beats every external baseline we scored on PublicSupported: 82.32% vs. the next-best baseline, PlantCLEF2024, at 82.14% (0.18 pp margin); the rest trail from TIPSv2-Giant at 76.1% down to DINOv2-LARGE at 64.58%, with TIPSv2-LARGE at 71.33% (Section 5)Public retrievalHigh over every baseline except PlantCLEF2024, which is too close to call
Retrieval capability beats every external baseline we scored on AgriStress-500 V1.0.2Supported: 73.95% vs. the next-best baseline, BioCLIP, at 73.05% (0.90 pp margin); the rest trail from PlantCLEF2024 at 70.43% down to DINOv2-LARGE at 54.77%, with TIPSv2-LARGE at 64.79% (Section 5)AgriStress-500 V1.0.2 retrievalHigh
Dense capability is the strongest of the five models at fine-grained crop-subtype granularitySupported: leads k-NN mIoU, k-NN accuracy, ARI and NMI against all four baselines; TIPSv2-base leads linear-probe mIoU, and the k-NN mIoU margin over TIPSv2-base sits inside the confidence intervalsFine-grained crop-subtype benchmark, 10 subtypes, five modelsHigh
Leads all four baselines on both weed labels that separateSupported: grass 0.570 and generic weed 0.904 k-NN mIoU, highest of the five on both, mean 0.737 vs. 0.720 / 0.684 / 0.677 / 0.615Fine-grained weed-subtype benchmark, 2 of 4 eligible labelsModerate
Greater local spatial continuity than DINOv2-LARGE and PlantCLEF2024Supported: 0.866 vs. 0.671 and 0.581 adjacent-patch similarity, below TIPSv2-LARGE (0.902)Current protocol; no boundary metricModerate
08

Intended & non-intended uses

Intended: retrieval capability

  • Whole-image semantic retrieval
  • Similar-image search
  • Candidate generation
  • Dataset exploration
  • Retrieval-based classification

Intended: dense representation capability

  • Sub-class semantic categorization and segmentation: separating individual crop and weed subtypes within a coarse class, and driving segmentation at that granularity. This is the primary intended use of the patch-token output and the capability with the strongest evidence in this report (Sections 6.1 and 6.2).
  • Frozen-feature dense classification, coarse and fine-grained
  • Spatial correspondence
  • Region retrieval
  • Localization
  • Input to downstream segmentation systems
Non-intended or unvalidated uses
  • Unsupervised segmentation without further validation
  • Pixel-accurate boundaries without a suitable decoder (this report's dense evaluation does not include boundary-level metrics)
  • Domains and imagery outside the tested suites
  • Cross-domain deployment based only on the benchmarks in this report
09

Limitations

  • We evaluate on the reported datasets only: two retrieval suites and one feature-separation benchmark.
  • Retrieval relevance is binary and category-driven; it does not capture graded or partial relevance.
  • Benchmark composition (difficulty, size, label source) differs across suites and affects comparability; results on one suite should not be assumed to generalize to the other.
  • Fine-grained weed-subtype results (Section 6.2) rest on two labels. The 0.737 mean IoU is arithmetic over the two per-label IoUs for grass and the generic weed class, not a separate benchmark run. Every other figure at this granularity is the benchmark's macro-average across all four eligible labels, including broadleaf and palmer amaranth, which score at or near zero IoU for every model and which we do not tabulate per label. Weed-species coverage is the limit here.
  • Adjacent-patch smoothness may reflect either useful spatial coherence or oversmoothing; we have not yet included a boundary-aware metric to distinguish the two.
  • We compute silhouette (Section 6.5) against ground-truth class labels, not cluster assignments; it does not move in the same direction as ARI/NMI in this revision's coarse-granularity data (Appendix B), and at fine-grained granularity is negative for most models.
  • Unsupervised clustering ranks the models differently by granularity. At fine-grained crop-subtype granularity AgriEmbedding has the highest ARI and NMI of the five models (0.386 / 0.557). At coarse granularity that reverses (DINOv2-LARGE highest, AgriEmbedding lowest of the five), and weed-subtype values stay modest in absolute terms for all five.
  • Dense evaluation in this report does not establish pixel-perfect segmentation capability.
  • Confidence intervals are only partially available. k-NN mIoU and linear-probe mIoU on the dense benchmark have bootstrapped 95% CIs for all five models, at coarse and fine-grained granularity alike (Appendix B). k-NN accuracy, ARI, NMI, silhouette, and adjacent-patch smoothness do not, nor does any retrieval metric; margins should be read as point estimates, not statistically distinguishable differences.
  • We document sample counts for both retrieval suites and the dense benchmark. Public retrieval covers 2,852 queries across 8 crop subsets, with a gallery equal to the query pool and relevant items per query equal to cluster size − 1 (Appendix C); AgriStress-500 V1.0.2 retrieval covers 452 queries across 122 field-plot subsets (39 with a single query), and we do not document its gallery size. The coarse dense benchmark covers 451 images and 449,982 patches.
  • Neither benchmark is size-matched to AgriEmbedding at 22M, though two retrieval baselines come close. Every dense-track baseline is at least 3.9x larger, and TIPSv2-LARGE and DINOv2-LARGE are roughly 14x larger. The retrieval track spans about 6M to 8B, and two of its baselines sit near our own size: AgriCLIP at ~25.6M is within about 16% of 22M, and EUPE-ViT-T at ~6M is the only model in this report smaller than the one under evaluation. Every other retrieval baseline is at least 3.9x larger, so read every result against the parameter counts in Section 5.
  • The dense benchmark covers five models, a smaller roster than the retrieval benchmark. The two are scored on different benchmarks against their own eligible baselines and we do not compare them against each other. Section 6 states the eligibility rule and gives our reason for each baseline it excludes: architectural for AgriCLIP, MaskDINO and the two Qwen sizes, and absence of a patch-token log for the other seven.
  • The dense benchmark is not field-disjoint. Reference and evaluation patches may come from the same field or image cluster, so some samples can have very similar neighbors from the same scene. The same split and protocol are used for all five models, so the relative results remain directly comparable, while the absolute scores are an upper bound, not a generalization estimate for a new field.
  • We include no qualitative examples in this revision: no retrieval successes or failures, no patch PCA visualizations, no prediction or boundary overlays.
  • We have not confirmed the exact MAP@10 denominator convention used by the underlying metrics library; low MAP@10 values on tasks with many valid matches per query should be read with that caveat.
  • We ran no robustness or sensitivity testing in this revision: nothing on resolution, compression, blur, lighting, gallery size, patch-grid resolution, occlusion, or choice of k. Findings apply to the imagery and conditions of the tested suites.
10

Conclusion

Within the benchmarks we ran, AgriEmbedding supports both capabilities we tested it on, though not uniformly. It leads every external baseline we scored on retrieval across both suites, effectively tied with PlantCLEF2024 on Public (Section 5). On the dense benchmark the answer depends on granularity: all four baselines beat it at the coarse 3-way split (Section 6.4), and it beats all four at fine-grained crop-subtype granularity on k-NN mIoU, k-NN accuracy, ARI and NMI, which is its clearest advantage in this report.

Fine-grained weed separation points the same direction but rests on the two labels, of four eligible, that separate at all for any model (Section 6.2). Neither result generalizes further than stated. Confidence intervals for most metrics, qualitative examples, and robustness testing (Section 9) are the main open items for the next revision.

References

Sources cited in the text.

  1. Steininger, D., Trondl, A., Croonen, G., Simon, J., and Widhalm, V. "The CropAndWeed Dataset: A Multi-Modal Learning Approach for Efficient Crop and Weed Manipulation." IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023. openaccess.thecvf.com/content/WACV2023/html/Steininger_The_CropAndWeed_Dataset_A_Multi-Modal_Learning_Approach_for_Efficient_Crop_WACV_2023_paper.html
  2. AIT Austrian Institute of Technology. CropAndWeed dataset and tooling (source repository). github.com/cropandweed/cropandweed-dataset
  3. Precision AI. AgriStress-500, version 1.0.2 (dataset card). Hugging Face, released under CC-BY-NC-4.0. huggingface.co/datasets/precisionaiinc/AgriStress-500

Appendices

+A · Metric definitions and formulas

nDCG@k: ratio of discounted cumulative gain to its ideal ordering:

Relevance is binary (0/1), so this reduces to a rank-discounted hit score: a correct match at rank 1 contributes more than the same match at rank 10.

Precision@k: fraction of the top-k retrieved results that are relevant (not "hit rate").

MAP@10: precision averaged across every rank a relevant item appears at, up to rank 10, then averaged across queries. On tasks with hundreds of valid matches per query, MAP@10 is capped low by construction, since it can only credit the ≤10 matches surfaced, a ceiling shared across every model scored on this benchmark. Exact denominator convention not confirmed (see Section 9).

k-NN mIoU: each held-out patch labeled by majority vote of nearest labeled reference patches in raw frozen feature space; scored per class with IoU, macro-averaged. No classifier trained.

Linear-probe mIoU: one linear classifier fit on frozen reference-patch features, scored the same way.

ARI / NMI: patches clustered with no label access; true labels used only to score cluster agreement, never to produce clusters.

Class-label silhouette: how cleanly each true class's patches separate geometrically. Ground-truth labels, not cluster assignments.

Adjacent-patch similarity ("smoothness"): mean cosine similarity between a patch and its immediate spatial neighbors.

+B · Full benchmark tables
coarse dense representation, AgriStress-500 V1.0.2
Modelk-NN mIoUk-NN acc.Linear-probe mIoUARINMISilhouetteSmoothness
AgriEmbedding57.91%80.90%56.13%0.0190.0180.0880.866
TIPSv2-base (baseline)61.16%81.20%60.04%0.1450.1900.0310.871
TIPSv2-LARGE (baseline)60.69%80.99%61.42%0.1640.1970.0230.902
DINOv2-LARGE (baseline)59.21%79.99%60.70%0.3600.2190.0250.671
PlantCLEF2024 (baseline)59.74%80.28%59.12%0.1270.1710.0530.581
95% confidence intervals (k-NN mIoU / linear-probe mIoU only)
Modelk-NN mIoU 95% CILinear-probe mIoU 95% CI
AgriEmbedding56.19–59.55%54.55–57.65%
TIPSv2-base59.51–62.70%58.26–61.62%
TIPSv2-LARGE58.89–62.30%59.59–63.05%
DINOv2-LARGE57.35–60.89%58.74–62.65%
PlantCLEF202457.70–61.47%57.23–60.81%
per-class k-NN mIoU / recall (coarse)
ModelIoU (soil)IoU (crop)IoU (weed)Recall (crop)Recall (weed)
AgriEmbedding0.7980.5700.3690.7020.437
TIPSv2-base0.7810.5840.4700.7190.590
TIPSv2-LARGE0.7780.5810.4620.7120.574
DINOv2-LARGE0.7720.5630.4410.7080.573
PlantCLEF20240.7720.5680.4520.6930.593

All five models separate soil most easily and weed least easily; weed is the sparsest class in the underlying imagery (Appendix C) and the hardest for every model here.

fine-grained crop subtypes (10 subtypes)
Modelk-NN mIoUk-NN acc.Linear-probe mIoUARINMISilhouette
AgriEmbedding76.64%86.68%79.01%0.3860.5570.089
TIPSv2-base (baseline)75.17%85.21%83.51%0.1550.281−0.015
TIPSv2-LARGE (baseline)69.61%81.52%83.43%0.1150.207−0.036
DINOv2-LARGE (baseline)55.77%70.41%75.02%0.0950.154−0.007
PlantCLEF2024 (baseline)64.43%76.89%73.16%0.0900.157−0.006
95% confidence intervals, fine-grained crop subtypes (k-NN mIoU / linear-probe mIoU only)
Modelk-NN mIoU 95% CILinear-probe mIoU 95% CI
AgriEmbedding70.39–81.51%72.79–83.75%
TIPSv2-base69.58–79.13%77.55–87.50%
TIPSv2-LARGE64.32–73.38%77.44–87.59%
DINOv2-LARGE51.54–58.26%69.49–78.37%
PlantCLEF202459.67–67.30%67.84–76.41%

Sample counts: 403,647 total patches (269,479 reference / 134,168 eval, 66.8%/33.2%).

fine-grained weed subtypes (macro-averaged over 4 eligible labels)
Modelk-NN mIoUk-NN acc.Linear-probe mIoUARINMISilhouette
AgriEmbedding21.30%90.67%21.20%0.0810.105−0.017
TIPSv2-base (baseline)21.43%89.84%21.41%0.0040.020−0.047
TIPSv2-LARGE (baseline)17.84%88.77%20.97%−0.0020.014−0.062
DINOv2-LARGE (baseline)15.63%87.45%23.83%0.0160.058−0.048
PlantCLEF2024 (baseline)19.89%88.97%20.89%−0.0030.012−0.025
95% confidence intervals, fine-grained weed subtypes (k-NN mIoU / linear-probe mIoU only)
Modelk-NN mIoU 95% CILinear-probe mIoU 95% CI
AgriEmbedding19.02–33.09%19.45–36.44%
TIPSv2-base18.86–38.85%18.97–39.06%
TIPSv2-LARGE15.84–32.12%18.30–32.80%
DINOv2-LARGE14.29–24.84%21.14–37.93%
PlantCLEF202418.08–35.82%18.56–32.79%

Sample counts: 226,259 total patches (150,738 reference / 75,521 eval, 66.6%/33.4%).

+C · Benchmark cards

Public retrieval suite. Built from the CropAndWeed dataset[1][2]. We build it from fine-grained bounding-box annotations with a fixed sampling seed, so construction is deterministic. This is CropAndWeed's "image-to-image" test set, one of three the dataset's own tooling generates (the other two are Plant-to-Plant and Plant-to-Image, below).

Public retrieval suite, construction
FieldValue
Crop categories8
Crop types (CropOrWeed2)Bean, Maize, Pea, Potato, Pumpkin, Soy, Sugar Beet, Sunflower
Cluster definitionOne dominant crop type per recording session, by bounding-box area
Cluster sizeMinimum 2, capped at 600 per crop type
Total images = total queries2,852
Relevant items per queryCluster size − 1 (variable, minimum 1)
GalleryThe same pool as the queries
Cap discrepancy

Two crop types exceed the documented 600-image cap in the query counts already reported in Appendix B: Maize (601) and Sugar Beet (606), over by one and six images respectively.

Other retrieval modes in this collection, not evaluated here: Plant-to-Plant (instance-crop-to-instance-crop, capped at 300 crops per category) and Plant-to-Image (a single instance crop retrieving full field images, e.g. 1,600 queries against a 4,691-image gallery for the crop-category test set).

retrieval query counts
FieldPublicAgriStress-500 V1.0.2
Total queries2,852452
Number of subsets8 (crop types)122 (field plots)
Subsets with only 1 query039
Gallery sizeSame as query count (2,852)Not documented
Two count discrepancies, not resolved in the source data

The retrieval harness reports 122 subsets (field plots) for AgriStress-500 V1.0.2, one more than the dataset card's 121 L2 clusters; and 452 total queries drawn from AgriStress-500 V1.0.2's 451 source images. We cannot explain either gap from the source data available for this revision, and neither is large enough to change the conclusions in Section 5.

We publish AgriStress-500 V1.0.2 under the precisionaiinc namespace, © Precision AI, under CC-BY-NC-4.0. That license also limits how readers of this report may reuse the benchmark commercially. Its published dataset card[3] reports the following.

AgriStress-500 V1.0.2, as published
FieldValue
Drone images from working fields500
Source images451
Segmentation masks451
Images with / without instance crops309 / 142
Plant instances4,041
Published dataset rows4,943
FormatPNG, variable resolution
Image width range96 px to ~6.02k px
Semantic classes present in masks18 of 29 defined
Crop types10
Weed-related classes8: Weed (generic), Grass, Broadleaf, Dandelion, Palmer Amaranth, Morning Glory, Redroot Pigweed, Lamb's-quarters
Dominant foreground areaWheat and Barley, ~66% combined
Clusters20 L1, 121 L2 sub-clusters
SplitsSingle train split
LicenseCC-BY-NC-4.0

Three of these figures carry caveats. The 4,943 row count is the Dataset Viewer's, and it does not correspond one-to-one with the image, mask, or instance counts; the row-level schema is not documented. Background appears in the imagery but is not one of the 18 classes. We do not document the resize or crop resolution our eval harness actually uses.

feature-separation benchmark split, resolved
FieldValue
Images used451
Total patches449,982
Reference / eval patches298,982 (66.4%) / 151,000 (33.6%)
Reference patches: soil / crop / weed178,182 / 79,536 / 41,264
Eval patches: soil / crop / weed90,992 / 39,998 / 20,010
Benchmark harness version2.0
+D · Reproducibility details

General controls: we take all embeddings from frozen models (no fine-tuning during eval), L2-normalize before comparison, and score AgriEmbedding and the external baselines through one identical pipeline per benchmark. We did not record random seeds, evaluation software version, or checkpoint IDs in the source data.

Retrieval protocol: every image/instance crop encoded once and cached → full similarity matrix per query, sorted into a ranked list → binary category-driven relevance (shared crop/weed/cluster/instance label) → Precision/nDCG/MAP computed per query, averaged. We do not document self-match removal, tie-handling, or CI methodology.

Patch-feature protocol: patch size 14, frozen features → reference/eval pool split → k-NN majority vote scored by IoU → linear probe (lbfgs, max 1,000 iter., converged for all five models in 101–448 iterations) fit on reference features, scored the same way → ARI/NMI/silhouette computed unsupervised, scored against true labels → smoothness = mean cosine similarity to grid neighbors.

Fine-grained evaluation scores individual crop and weed subtypes directly rather than the coarse buckets. We only score a candidate subtype ("eligible") if it clears a minimum-sample threshold in both pools: all 10 crop subtypes clear it, while only 3 named weed species plus a generic catch-all do.

The dense benchmark and the AgriStress-500 retrieval suite both run on the public AgriStress-500 V1.0.2 test set[3]; the Public retrieval suite is built from CropAndWeed[1] (Appendix C).

+E · Terminology
AvoidedUsed instead
"Hit rate" for Precision@k"Fraction of the top-k results that are relevant"
"The advantage is shrinking""The advantage varies by benchmark"
"Smoothness confirms specialization""Consistent with greater local continuity"
"Segmentation quality" from silhouette alone"Cluster geometry" / "class separability"
"Patch wins segmentation""AgriEmbedding leads frozen fine-grained dense-prediction metrics"
"CLS variant" / "Patch variant" (implying two separate models)"AgriEmbedding" (one model), described by which output is in use