Today we are releasing Manta, a visual evidence model for work that gets questioned afterwards. Calling it a geolocation model undersells it: locating the photograph is where the work starts.
Manta establishes where an image was taken from its content alone, with no metadata. It then resolves that answer down to the street or the building rather than stopping at a radius. It returns the evidence that led there, so the conclusion can be argued with. And it is held to that evidence. The precision it states limits what it is allowed to claim, so when the evidence does not support a street, it does not name one.
Manta is one AI, doing two things. Geo-Estimation establishes where the photograph was taken. Street-Match resolves that to the street or the building. It is one model, and a single evidence chain runs from the first answer to the last.
This is what we mean by physical inference: a fact about the physical world, established from what the camera recorded, with the evidence attached and the claim limited to what that evidence supports. Manta is the part of it we can prove today, and the rest of this post is that proof.
Manta joins Orca rather than replacing it. Orca continues to handle everyday location work. Manta is for the cases where somebody will disagree with the answer: a disputed claim, a compliance record, an editorial decision, a file that another person reads months later and has to trust.
What changed
Earlier models returned a location and a score. Nothing tied the score to the answer, so a model could name a specific street corner while reporting that it was only sixty percent sure. Both cannot be true, which leaves the score as decoration.
In Manta, precision is bound to evidence. When the visual evidence supports a region, it names a region. When it genuinely resolves to a single site, it says so and the confidence rises to match. The model no longer produces answers more specific than it can defend, which is a narrower behaviour than before and a more useful one.
The second change is what comes back with the answer. Manta returns the observations it made, the inferences drawn from each, the candidate locations it considered and set aside, and why. A model that shows only its best answer is asking to be trusted. One that shows its alternatives is giving you something to check.
Reading the picture
Manta works the way a careful analyst does, without getting bored. Which side of the road the traffic is on. The species in the hedge and whether it survives a winter. The shape of the utility hardware, the font on the signage, the pitch of the roofs, the quality of the light and what latitude produces it. None of these are conclusive alone. Together they narrow the world down.
A record, not just a result
The original image is hashed on arrival, so the file that was analysed can be checked against the copy anyone else holds. Around that anchor, every Manta analysis produces a record built to be read by somebody who was not there: the evidence chain, the rejected candidates, the confidence and its basis, all captured and exported as a single document. This is the part that matters when the question stops being what happened and becomes how do you know.
Benchmarks
Geo-estimation’s latest OSV-5M board reports 45.77% within 25 km, 81.83% within 200 km, 92.57% within 750 km, and 97.30% within 2,500 km, with a median error of 30.5 km. Each is ahead of Chipoint-2’s self-reported model card: 44.83%, 81.10%, 92.11%, and 96.82%, and a 32 km median[30].
The margins are +0.94, +0.73, +0.46, and +0.48 percentage points, and 1.5 km lower on median error. These are numerical leads; the smaller differences should not be read as established improvements beyond evaluation variability.
| OSV-5M | Geo-estimationOceanir | Chipoint-2[30]self-reported | Chipoint[2]v1, self-reported | Pinpoint[3]paper, closed weights | GeoCLIP[4]published | Lead |
|---|---|---|---|---|---|---|
| Within 25 kmCity | 45.77 | 44.83 | 37.63 | 35.60 | 21.50 | +0.94 |
| Within 200 kmRegion | 81.83 | 81.10 | 70.56 | 67.50 | 52.10 | +0.73 |
| Within 750 kmCountry | 92.57 | 92.11 | 84.94 | 83.70 | 72.10 | +0.46 |
| Within 2,500 kmContinent | 97.30 | 96.82 | 93.36 | 93.20 | — | +0.48 |
| Median errorLower is better | 30.5 km | 32 km | 51 km | — | — | −1.5 km |
Im2GPS3k, YFCC4k and GWS15k are classical geolocation splits the field has measured itself against for a decade. Our runs on all three are reported below, alongside YFCC26k. The bars quoted are the strongest published numbers we could verify against the source paper, not leaderboard scrapes; all three are reported at five thresholds in the literature (1 km, 25 km, 200 km, 750 km, 2,500 km) and 25 km is the city-level figure shown here. No single method holds all three, which is the point. Manta now leads YFCC4k and Im2GPS3k at all five thresholds. On GWS15k it leads LocDiff[34], PIGEOTTO[28] and GeoSURGE[29] at every threshold, compared against PIGEOTTO’s headline result. The narrowest lead is at 1 km, 2.39 against LocDiff’s 2.1. On Im2GPS3k Manta takes every threshold. The 1 km lead over REVERSE[24] is 22.52 against 22.5, too small to call a win; the 2,500 km lead over Pinpoint[3] is 91.16 against 90.2. One method is listed on that board but not ranked: GeoRouter[32], at 50.48 at 25 km against our 50.82. It is not a geolocation model but a routing layer over a general-purpose model, which its own table reports at 47.91 before routing. That is a statement about method class and not about its numbers, which we have no reason to doubt. GeoSearch[33] is not listed at all. It runs a web-scale reverse image search and reads the pages it finds, so its score measures what the web already knows about a photograph as much as what can be inferred from it. Its authors evaluate with leakage in mind, and we do not dispute their numbers; they answer a different question. Photo-benchmark rows, ours included, can overlap with the training data the field shares; we flag it rather than claim immunity. GWS15k is by far the hardest because it samples Street View near random city centres worldwide, so 92 percent of its locations never appear in the training corpus the field shares. MP-16[27] is excluded because it is that shared training corpus rather than a test set: no standard held-out split, so no bar to clear.
Geolocation across benchmarks
On the closed-system vendors
The obvious question is how this compares to the closed-system tools an agency already evaluates. The short answer is that it cannot be compared, and we would rather say so plainly than work around it.
A closed-system vendor publishes a worked example in place of an evaluation. It shows one photograph and one result: a radius or an address and a handful of ranked candidates. There is no dataset, sample size, accuracy rate or standard split behind it. A worked example is a screenshot of one good day, and there is nothing in it a buyer can check.
The same pages carry a second thing, further down and in smaller type: that the output is an investigative lead rather than a confirmed location, and that the similarity score ranks candidates without being a probability that any one of them is right. Those caveats are correct, and we would write them ourselves. They are why confidence here limits what a result is allowed to claim instead of sitting beside it as decoration.
But a closed system cannot hold both ends. A page cannot advertise meter-level resolution at the top and disclaim at the bottom that its ranking carries no probability. One of those is the product and the other is the legal position, and an agency finds out which one it bought after it has signed. That is what closed means in practice. Private weights are ordinary; private evidence is the problem.
So there is no closed-system column in the tables above. A worked example cannot be placed beside 21,131 images without flattering somebody dishonestly, and we would rather not do that in either direction. What can be compared is how each side reports. Every number here sits on a public split, at a stated sample size, against named methods with citations, including the splits where Manta does not lead. That is a claim an agency can audit before it signs, and it is the only difference that does not require taking a vendor at their word.
The full results go to organizations evaluating Manta, including the splits where it does not lead. The report carries the OSV-5M full-test row and where it sits against the current bar, every classical split with the competing methods beside it, the failure modes we know about, and the evaluation protocol we recommend for testing on your own material.
Request the technical report · Request organization access
YFCC4k update
Geo-Estimation’s new YFCC4k run is ahead at all five thresholds. The best-other comparator is selected per threshold, so the margin is against a different method at different distances.
| Threshold | Manta | Best other | Lead |
|---|---|---|---|
| @1 km | 35.04 | 33.00 (Pinpoint) | +2.04 |
| @25 km | 47.39 | 44.40 (Pinpoint) | +2.99 |
| @200 km | 58.50 | 57.50 (Pinpoint) | +1.00 |
| @750 km | 73.15 | 71.80 (Pinpoint) | +1.35 |
| @2500 km | 84.56 | 84.50 (Pinpoint) | +0.06 |
The OSV-5M table uses Oceanir’s latest reported board and highlights the highest numerical score at each radius. Run-to-run variability matters when interpreting small margins. The median error comes from the same predictions, measured by a standalone scorer; Chipoint-2 reports its median rounded to the whole kilometre. Lead is Manta minus the comparison score; positive is better. Pinpoint’s YFCC4k curve is taken from its paper, which publishes every threshold.
| Method | @1 km | @25 km | @200 km | @750 km | @2500 km |
|---|---|---|---|---|---|
Im2GPS3k[26] 2,997 images. The split most papers lead with. | |||||
| Geo-Estimation (Manta)bar | 22.52 | 50.82 | 66.57 | 80.91 | 91.16 |
| REVERSE[24] | 22.5 | 48.3 | 59.3 | 73.5 | 84.8 |
| Pinpoint[3] | 20.5 | 47.4 | 63.5 | 79.0 | 90.2 |
| Geo-ADAPT[37] | 17.9 | 45.3 | 62.6 | 77.9 | 89.5 |
| GeoRanker[25] | 18.79 | 45.05 | 61.49 | 76.31 | 89.29 |
| HierLoc[39] | 11.3 | 43.8 | 58.4 | 74.1 | 85.1 |
| GeoSURGE[29] | 17.2 | 42.5 | 58.1 | 74.6 | 87.6 |
| GeoMetric[38] | 17.7 | 42.2 | 56.8 | 71.7 | 86.1 |
| Geo-R[43] | 18.10 | 41.53 | 58.31 | 75.33 | 86.42 |
| DualGeo[35] | 17.25 | 41.47 | 55.76 | 71.71 | 85.05 |
| Chipoint v2[30] | 16.5 | 41.2 | 54.1 | 70.7 | 84.8 |
| G3[25] | 16.65 | 40.94 | 55.56 | 71.24 | 84.68 |
| GLOBE[44] | 9.84 | 40.18 | 56.19 | 71.45 | 82.38 |
| Img2Loc[25] | 15.34 | 39.83 | 53.59 | 69.70 | 82.78 |
| GeoToken[36] | 16.8 | 39.6 | 53.8 | 70.8 | 85.0 |
| GaGA (MP16-Pro)[42] | 15.0 | 37.1 | 49.5 | 67.3 | 82.4 |
| PIGEON[28] | 11.3 | 36.7 | 53.8 | 72.4 | 85.3 |
| LocDiff-H[34] | 15.3 | 36.5 | 56.4 | 75.2 | 87.4 |
| GRE Suite[31] | 11.30 | 35.30 | 51.70 | 69.30 | 85.70 |
| GeoCLIP[4] | 14.11 | 34.47 | 50.65 | 69.67 | 83.82 |
| Concept-Aware[46] | 13.2 | 34.0 | 49.8 | 68.2 | 83.5 |
| GaGA[42] | 11.7 | 33.0 | 48.0 | 67.1 | 82.1 |
| GeoLocSFT[45] | 8.80 | 32.70 | 47.20 | 65.80 | 82.25 |
| StreetCLIP[48] | — | 22.4 | 37.4 | 61.3 | 80.4 |
Pipelines over general-purpose models A routing layer, tool loop or prompting stage wrapped around a general-purpose model such as Gemini or GPT-4V, rather than geolocation weights of its own. Shown in full and not ranked above, because each row measures the wrapper as much as the model. Their numbers are not in dispute. | |||||
| GeoRouter[32] | 20.82 | 50.48 | 65.73 | 80.35 | 90.66 |
| GeoToken + Gemini[36] | 19.0 | 46.0 | 60.1 | 76.6 | 88.8 |
| VLM-VPR[41] | 18.62 | 45.65 | 59.79 | 73.71 | 85.95 |
| SpotAgent[40] | 14.12 | 40.36 | 57.80 | 73.43 | 85.75 |
| NAVIG[47] | 5.5 | 28.9 | 49.1 | 68.3 | 84.0 |
YFCC4k[26] 4,536 Flickr images. Noisier, more indoor and close-up. | |||||
| Geo-Estimation (Manta)bar | 35.04 | 47.39 | 58.50 | 73.15 | 84.56 |
| Pinpoint[3] | 33.0 | 44.4 | 57.5 | 71.8 | 84.5 |
| GeoRanker[25] | 32.94 | 43.54 | 54.32 | 69.79 | 82.45 |
| GeoMetric[38] | 31.2 | 41.1 | 51.8 | 67.8 | 81.0 |
| Geo-ADAPT[37] | 32.5 | 39.1 | 55.4 | 70.8 | 84.5 |
| DualGeo[35] | 27.49 | 36.45 | 45.03 | 61.58 | 75.92 |
| G3[25] | 23.99 | 35.89 | 46.98 | 64.26 | 78.15 |
| GeoToken[36] | 24.3 | 35.3 | 46.6 | 64.2 | 78.6 |
| Chipoint v2[30] | 19.1 | 34.5 | 44.2 | 60.2 | 76.0 |
| GaGA (MP16-Pro)[42] | 24.3 | 33.7 | 43.9 | 61.4 | 76.3 |
| GeoSURGE[29] | 19.9 | 33.6 | 48.7 | 67.4 | 82.0 |
| RFM S2 (10M)[49] | — | 33.5 | 45.3 | 61.1 | 77.7 |
| Img2Loc[25] | 19.78 | 30.71 | 41.40 | 58.11 | 74.07 |
| HierLoc[39] | 8.4 | 30.2 | 43.3 | 61.7 | 75.8 |
| REVERSE[24] | 14.1 | 27.5 | 38.1 | 53.8 | 70.6 |
| PIGEON[28] | 10.4 | 23.7 | 40.6 | 62.2 | 77.7 |
| RFM S2[49] | — | 23.7 | 36.4 | 54.5 | 73.6 |
| Geo-R[43] | 10.47 | 22.67 | 40.04 | 60.83 | 75.84 |
| GaGA[42] | 6.9 | 18.9 | 34.5 | 56.7 | 71.6 |
| GeoLocSFT[45] | 5.21 | 18.58 | 32.64 | 53.10 | 72.10 |
Single-threshold reports Published at one distance only, so they cannot be ranked against the curves above. | |||||
| GeoCLIP[4] | 10.0 | — | — | — | — |
Pipelines over general-purpose models A routing layer, tool loop or prompting stage wrapped around a general-purpose model such as Gemini or GPT-4V, rather than geolocation weights of its own. Shown in full and not ranked above, because each row measures the wrapper as much as the model. Their numbers are not in dispute. | |||||
| GeoRouter[32] | 32.98 | 46.01 | 57.52 | 72.02 | 83.02 |
| GeoToken + Gemini[36] | 25.4 | 38.5 | 51.4 | 68.0 | 81.0 |
| SpotAgent[40] | 7.30 | 21.52 | 36.18 | 55.00 | 70.77 |
YFCC26k[26] 21,131 Flickr images. The larger YFCC split. Competitor curves are taken from GeoSURGE Table 4, which reproduces every prior method on this split under one protocol, except PIGEOTTO, which is shown at its own paper's headline result rather than the weaker ME16-only variant. Read the margin here with that in mind: the strongest recent retrieval and LVLM methods have not published on 26k. Pinpoint, GeoRanker, GeoRouter, REVERSE, G3, DualGeo and Chipoint v2 all report on YFCC4k, where the top of the field sits within about two points of us at 1 km, and none of them report 26k. LocDiff does, in its hybrid form, and is listed. So the field on this split is LocDiff, GeoSURGE and older work, and the double-digit gap is partly a statement about who has been measured rather than how far ahead we are. We would rather say that than let the two YFCC tables be lined up against each other and have it said for us. | |||||
| Geo-Estimation (Manta)bar | 30.68 | 45.31 | 58.05 | 73.48 | 86.02 |
| GeoMetric[38] | 20.3 | 33.7 | 46.4 | 64.1 | 79.3 |
| GeoSURGE[29] | 17.8 | 31.5 | 45.1 | 64.3 | 79.3 |
| RFM (S2-10M)[29] | 5.3 | 29.0 | 40.9 | 57.8 | 75.8 |
| LocDiff-H[34] | 13.2 | 26.0 | 41.9 | 64.5 | 80.3 |
| PIGEOTTO[28] | 10.5 | 25.8 | 42.7 | 63.2 | 79.0 |
| GeoDecoder[28] | 10.1 | 23.9 | 34.1 | 49.6 | 69.0 |
| GeoCLIP[4] | 11.6 | 22.2 | 36.7 | 57.5 | 76.0 |
| TransLocator[28] | 7.2 | 17.8 | 28.0 | 41.3 | 60.6 |
| GR/Qwen-VL[29] | 4.0 | 17.4 | 28.9 | 48.1 | 67.8 |
| ISNs[28] | 5.3 | 12.3 | 19.0 | 31.9 | 50.7 |
| PlaNet[29] | 4.4 | 11.0 | 16.9 | 28.5 | 47.7 |
GWS15k[28] Street View shots near random city centres. The released split resolves to 14,955 images; Manta is scored against all of them, with the 394 that could not be scored counted as misses. 92 percent of its locations are unseen in the training corpus the field shares, which is why every number here collapses. Manta leads every method at all five thresholds, including LocDiff, the strongest published result below 2,500 km: 1 km (2.39 against LocDiff's 2.1), 25 km (15.23 against 12.4), 200 km (45.67 against 33.7), 750 km (78.00 against 67.0) and 2500 km (93.62 against PIGEOTTO's 85.1). The 1 km lead is 0.29 points, the narrowest of the five. | |||||
| Geo-Estimation (Manta)bar | 2.39 | 15.23 | 45.67 | 78.00 | 93.62 |
| LocDiff[34] | 2.1 | 12.4 | 33.7 | 67.0 | 85.0 |
| PIGEOTTO[28] | 0.7 | 9.2 | 31.2 | 65.7 | 85.1 |
| GeoSURGE[29] | 1.0 | 4.6 | 21.9 | 54.7 | 80.8 |
| GRE Suite[31] | 0.9 | 4.1 | 18.9 | 54.8 | 78.3 |
| GeoCLIP[4] | 0.6 | 3.1 | 16.9 | 45.7 | 74.1 |
| GeoDecoder[28] | 0.7 | 1.5 | 8.7 | 26.9 | 50.5 |
| TransLocator[28] | 0.5 | 1.1 | 8.0 | 25.5 | 48.3 |
| ISNs[28] | 0.05 | 0.6 | 4.2 | 15.5 | 38.5 |
Pipelines over general-purpose models A routing layer, tool loop or prompting stage wrapped around a general-purpose model such as Gemini or GPT-4V, rather than geolocation weights of its own. Shown in full and not ranked above, because each row measures the wrapper as much as the model. Their numbers are not in dispute. | |||||
| VLM-VPR[41] | 0.48 | 10.28 | 32.61 | 64.14 | 85.25 |
Accuracy at each distance threshold, self-reported by each method unless noted. Groups are ordered by how many of the five thresholds each row tops, ties resolved at 25 km, the same measure the bar uses. “Bar” is derived, not assigned: within each split it marks whichever full-curve method tops the most of the five thresholds, with ties resolved at 25 km. Methods that publish a single threshold are listed separately and never badged, because a method cannot win a comparison it only partly enters. Rows are not perfectly like-for-like: methods differ in how they are built and trained. Nor are they equally checkable. Pinpoint posts the highest non-Manta figure on YFCC4k at 1 km (33.0). Its figures come first-hand from Virtualitics’ paper[3], which reports every threshold we quote, and its code is released for academic analysis. The trained weights are not, a choice the paper ties to misuse risk, so its numbers can be read at their source but not re-run by us or anyone else. We report them anyway, because leaving out the closest result to our own would be the more convenient omission, but a number you can read is not the same kind of evidence as a number you can re-run, and the row does not say so on its own. No method leads all four classical splits.
VPR benchmarks
Street-Match is the mode that resolves a photograph to the exact street or building. These are Manta’s results on seven standard VPR benchmarks, alongside every published model we could find numbers for. Recall@1 throughout.
What this table compares. Every row is a model: a specific set of trained weights evaluated end to end. It is not a comparison of techniques. Aggregation schemes, mining strategies and fine-tuning recipes are how these models are built, and several rows here share them, but a technique has no Recall@1 and cannot be put in a column. Where a paper reports one idea in several configurations we take its headline model rather than listing the idea once and the ablations beside it.
And where we sit in it. A ◇ marks a published result with no public code or weights. Street-Match carries one too. It is proprietary, we do not release its weights, and nobody outside Oceanir can reproduce our rows from scratch — so we are in exactly the category we are flagging, and it would be dishonest to mark competitors for it and exempt ourselves. What we can offer instead of weights is your own imagery: send the hardest cases you have and we will return what the system produces. Rows without a ◇ you can download and run today, and the citation links the weights. We keep the ◇ rows rather than dropping them, because two of them beat us and removing a competitor that wins is the more self-serving edit, not the more rigorous one.
A dash means the model was not evaluated on that benchmark in its paper or in subsequent reproductions.
| Model | Year | Reported | SPED | Pitts30k | MSLS-val | Nordland | AmsterTime | Tokyo 24/7 | SVOX |
|---|---|---|---|---|---|---|---|---|---|
| 1SAGE (8448-D)[18] | 2026 | 6/7 | 98.9 | 95.8 | 94.5 | 96.0 | 83.5 | 97.5 | — |
| 2Street-Match 1.0 ◇ | 2026 | 7/7 | 95.39 | 94.95 | 94.60 | 97.21 | 74.25 | 98.43 | 98.74 |
| 3SAGE-L (no InteractHead)[18] | 2026 | 7/7 | 92.1 | 94.7 | 94.2 | 94.8 | 65.6 | 98.1 | 98.8 |
| 4BoQ[12] | 2024 | 7/7 | 92.5 | 93.7 | 93.8 | 90.6 | 63.0 | 98.1 | 99.0 |
| 5SuperPlace NVL-FT² (B) ◇[16] | 2025 | 7/7 | 87.5 | 93.7 | 94.3 | 91.4 | 62.3 | 96.8 | 98.6 |
| 6EMVP ◇[14] | 2024 | 6/7 | 94.6 | 94.0 | 93.9 | 88.7 | 65.6 | 96.8 | — |
| 7SuperVLAD[13] | 2024 | 6/7 | 93.2 | 95.0 | 92.2 | 91.0 | 63.9 | 95.6 | — |
| 8SALAD[11] | 2024 | 7/7 | 92.1 | 92.4 | 92.2 | 90.0 | 58.8 | 94.6 | 98.2 |
| 9FoL[17] | 2025 | 6/7 | 92.1 | 93.9 | 93.1 | 87.8 | 64.6 | 96.2 | — |
| 10SALAD-CM[11] | 2024 | 6/7 | 89.5 | 92.6 | 94.2 | 95.6 | 57.8 | 96.8 | — |
| 11SelaVPR †[9] | 2024 | 7/7 | 88.6 | 92.8 | 90.8 | 87.3 | 55.2 | 94.0 | 97.2 |
| CricaVPR[10] | 2024 | 5/7 | 91.3 | 94.9 | 90.0 | 90.7 | — | 93.0 | — |
| MixVPR[7] | 2023 | 5/7 | 84.7 | 91.5 | 88.0 | 76.2 | — | 85.1 | — |
| EigenPlaces[8] | 2023 | 5/7 | 70.2 | 92.5 | 89.1 | 71.2 | — | 93.0 | — |
| CosPlace[6] | 2022 | 5/7 | 75.5 | 88.4 | 82.8 | 58.5 | — | 87.3 | — |
| NetVLAD[5] | 2016 | 5/7 | 70.2 | 81.9 | 53.1 | 6.4 | — | 60.6 | — |
| SuperPlace NVL-FT² (L) ◇[16] | 2025 | 3/7 | — | 94.1 | 94.5 | — | — | 97.1 | — |
| EffoVPR † ◇[15] | 2025 | 3/7 | — | 93.9 | 92.8 | — | — | 98.7 | — |
All values are Recall@1. Our own rows are set in italics throughout this page: they are Oceanir runs, not figures reported by a third party. Street-Match 1.0 is evaluated one query image at a time on all seven benchmarks. † indicates a two-stage method with re-ranking. Published numbers sourced from the original papers and subsequent reproductions (SAGE, ICLR 2026; SuperPlace, 2025; BoQ, CVPR 2024). Where a paper reports several configurations we quote its headline one rather than mixing rows. SAGE is listed twice because the paper reports it two ways. Its InteractHead applies attention across a batch while building the training graph, and the paper states that this cost sits in training and leaves inference unaffected, so both rows are evaluated per image at inference, the same way Street-Match is evaluated. The second row is the paper’s Table 9 ablation, trained without that head. We show both rather than whichever suits us. Nordland uses the full summer-vs-winter protocol. MSLS-val uses 25 m with azimuth within 40 degrees. Manta’s Nordland evaluation differs slightly in frame alignment. MSLS-val is scored over the published 740 queries. SVOX protocol is unverified. AmsterTime is the cross-era archival-to-modern matching task that our product solves daily, and the benchmark we care about most.
Seven benchmarks in, seven reported, Street-Match 1.0 ranks second of the eleven models that report at least six. We are saying second because second is what the numbers say. The model above us is SAGE at 8448 dimensions[18]: peer reviewed at ICLR 2026, weights released under MIT, evaluated per image the same way we are, and it beats us on five of seven boards. A shipping product placing second to the current state of the art, on every benchmark the field uses, is a result we are content to publish unedited. First place on a table we drew ourselves would be worth less.
Nordland is the one board where we lead SAGE, and the only margin in this row wide enough to mean anything.97.21 against its 96.0, and against SALAD-CM’s 95.6[11], over 27,592 queries, where the 95 percent band is 0.21 points. A lead six times its own noise floor on the largest split in the table is a real result, and it is a lead over the model at the top of the table rather than over the field with that model set aside.
The rest of the row is read against the same test the OSV-5M table uses: a margin counts only when it clears the sampling band for that split’s query count. Three boards are inside it and we record them as ties. MSLS-val[20] is 94.60 against SAGE’s 94.5, but 740 queries put the band at 1.64 points, so a tenth of a point decides nothing. Tokyo 24/7[22] is 98.43 against EffoVPR’s 98.7[15] on 315 queries, and SVOX is 98.74 against BoQ’s 99.0[12] on 823. Small splits cannot separate the top of this table from itself.
Three are separable losses and we are not going to round them off. Pitts30k[23] is 94.95, 0.85 behind SAGE. SPED is 95.39, 3.51 behind. AmsterTime[19] is 74.25, clear of every other model in the table but 9.25 behind SAGE’s 83.5. AmsterTime is the cross-era archival-to-modern task our product does every day, so the board we care about most is the one we trail by the widest margin.
Against the full field that is one win, three ties and three losses. SAGE leads five of the seven boards and the table overall, 94.37 against our 93.37. More on the work in the Street-Match post.
Limits
Interiors give it very little. Overcast skies remove the light it reads latitude from. A stretch of motorway in one temperate country looks a great deal like a stretch of motorway in another, and there are photographs where the honest answer is a region and a shrug. We would rather say so here than have you find out in front of somebody who was relying on the number.
Manta answers where a photograph was taken. It is not built to find people and we do not permit it to be used that way. That is written into our terms rather than left to good intentions. The question worth answering is whether an image is what it claims to be, and answering it does not require pointing anything at anybody.
The Manta suite
Manta is a single model. It does two jobs, each a mode that answers a different question and fails differently. Geo-estimation says roughly where on Earth. Street-Match says this exact place. You choose the mode that fits the work.
Street-Match resolves a photograph to the exact street or building. The VPR benchmark results above are Manta’s, running in Street-Match mode. It is in research preview now.
How it ships
Manta is a different kind of system than anything we have shipped before, built to commit to a claim and hold ground on it rather than offer a plausible answer and move on. A model that behaves that way earns its rollout in stages.
Enterprise organizations first, then selected pilots and research partners, then wider access. Controlled access is a stage, not a destination. For now, access is granted to an organization, not a person: we verify the organization, it names the people who will use it and what they will use it for, and the grant is scoped to that. Coverage is enabled per city against a declared purpose rather than switched on globally, because “which cities can you do” and “what are you doing with them” are the same question. If a request cannot survive being written down, it does not get provisioned.
We are starting with organizations that can put Manta inside a reviewed workflow, a compliance record, an investigation, a case file with someone accountable for what it concludes. We do not yet know how it behaves across the full range of images people submit outside that setting, and we would rather learn on our own terms than find out from the outside.
If the capability gives you pause, you have read it correctly. It gives us pause too. That is why it ships this way, and it is no reason to keep quiet about it.
Availability
Manta is available now to select organizations under a reviewed enterprise workflow. Orca remains the default across the web app and the API for everyone else, and existing integrations continue to work unchanged. Individuals keep Orca, which answers where a photograph was taken and stops there.
What we learn from Manta is what moves into Orca, so the improvement reaches everyone without carrying over the parts that are still unproven.
Manta is not self-serve, and the web app runs Orca. If you want to see how Manta handles your material, send the hardest cases you have, including ones you already know the answer to. Request organization access · [email protected]
Go deeper: Street-Match · Newsroom case study