Oceanir
  • Pricing
Book a demoAnalyze an image free
September 27, 2026/22 min·Announcements·Research·Benchmarks

Introducing Manta

A visual evidence model for work that gets questioned. Manta establishes where a photograph was taken, resolves it to the street or the building, shows the evidence that led there, and never claims more precision than that evidence carries.

Oceanir Team
A manta ray drawn in fine white contour lines on black, wings swept wide and tail trailing below.
Higher is better
  1. 01Geo-estimation45.67%
  2. 02LocDiff33.70%
  3. 03VLM-VPR32.61%
  4. 04PIGEOTTO31.20%
  5. 05GeoSURGE21.90%
  6. 06GRE Suite18.90%
  7. 07GeoCLIP16.90%
  8. 08GeoDecoder8.70%
  9. 09TransLocator8.00%
  10. 10ISNs4.20%
0%25%50%

Today we are releasing Manta, a visual evidence model for work that gets questioned afterwards. Calling it a geolocation model undersells it: locating the photograph is where the work starts.

Manta establishes where an image was taken from its content alone, with no metadata. It then resolves that answer down to the street or the building rather than stopping at a radius. It returns the evidence that led there, so the conclusion can be argued with. And it is held to that evidence. The precision it states limits what it is allowed to claim, so when the evidence does not support a street, it does not name one.

Manta is one AI, doing two things. Geo-Estimation establishes where the photograph was taken. Street-Match resolves that to the street or the building. It is one model, and a single evidence chain runs from the first answer to the last.

Manta · Demo. One photo, from upload to the matched street.

This is what we mean by physical inference: a fact about the physical world, established from what the camera recorded, with the evidence attached and the claim limited to what that evidence supports. Manta is the part of it we can prove today, and the rest of this post is that proof.

Manta joins Orca rather than replacing it. Orca continues to handle everyday location work. Manta is for the cases where somebody will disagree with the answer: a disputed claim, a compliance record, an editorial decision, a file that another person reads months later and has to trust.

Manta · The film

What changed

Earlier models returned a location and a score. Nothing tied the score to the answer, so a model could name a specific street corner while reporting that it was only sixty percent sure. Both cannot be true, which leaves the score as decoration.

In Manta, precision is bound to evidence. When the visual evidence supports a region, it names a region. When it genuinely resolves to a single site, it says so and the confidence rises to match. The model no longer produces answers more specific than it can defend, which is a narrower behaviour than before and a more useful one.

The second change is what comes back with the answer. Manta returns the observations it made, the inferences drawn from each, the candidate locations it considered and set aside, and why. A model that shows only its best answer is asking to be trusted. One that shows its alternatives is giving you something to check.

Reading the picture

Manta works the way a careful analyst does, without getting bored. Which side of the road the traffic is on. The species in the hedge and whether it survives a winter. The shape of the utility hardware, the font on the signage, the pitch of the roofs, the quality of the light and what latitude produces it. None of these are conclusive alone. Together they narrow the world down.

A record, not just a result

The original image is hashed on arrival, so the file that was analysed can be checked against the copy anyone else holds. Around that anchor, every Manta analysis produces a record built to be read by somebody who was not there: the evidence chain, the rejected candidates, the confidence and its basis, all captured and exported as a single document. This is the part that matters when the question stops being what happened and becomes how do you know.

Benchmarks

Geo-estimation’s latest OSV-5M board reports 45.77% within 25 km, 81.83% within 200 km, 92.57% within 750 km, and 97.30% within 2,500 km, with a median error of 30.5 km. Each is ahead of Chipoint-2’s self-reported model card: 44.83%, 81.10%, 92.11%, and 96.82%, and a 32 km median[30].

The margins are +0.94, +0.73, +0.46, and +0.48 percentage points, and 1.5 km lower on median error. These are numerical leads; the smaller differences should not be read as established improvements beyond evaluation variability.

OSV-5MGeo-estimationOceanirChipoint-2[30]self-reportedChipoint[2]v1, self-reportedPinpoint[3]paper, closed weightsGeoCLIP[4]publishedLead
Within 25 kmCity45.7744.8337.6335.6021.50+0.94
Within 200 kmRegion81.8381.1070.5667.5052.10+0.73
Within 750 kmCountry92.5792.1184.9483.7072.10+0.46
Within 2,500 kmContinent97.3096.8293.3693.20—+0.48
Median errorLower is better30.5 km32 km51 km——−1.5 km
Share of OSV-5M test photographs placed within each distance of where they were taken; higher is better, except median error. Bars are scaled to the best result in each row. Lead is Geo-estimation against the best other result. Chipoint-2 and Chipoint are self-reported; Pinpoint is from its own paper, and its trained weights are not public.

Im2GPS3k, YFCC4k and GWS15k are classical geolocation splits the field has measured itself against for a decade. Our runs on all three are reported below, alongside YFCC26k. The bars quoted are the strongest published numbers we could verify against the source paper, not leaderboard scrapes; all three are reported at five thresholds in the literature (1 km, 25 km, 200 km, 750 km, 2,500 km) and 25 km is the city-level figure shown here. No single method holds all three, which is the point. Manta now leads YFCC4k and Im2GPS3k at all five thresholds. On GWS15k it leads LocDiff[34], PIGEOTTO[28] and GeoSURGE[29] at every threshold, compared against PIGEOTTO’s headline result. The narrowest lead is at 1 km, 2.39 against LocDiff’s 2.1. On Im2GPS3k Manta takes every threshold. The 1 km lead over REVERSE[24] is 22.52 against 22.5, too small to call a win; the 2,500 km lead over Pinpoint[3] is 91.16 against 90.2. One method is listed on that board but not ranked: GeoRouter[32], at 50.48 at 25 km against our 50.82. It is not a geolocation model but a routing layer over a general-purpose model, which its own table reports at 47.91 before routing. That is a statement about method class and not about its numbers, which we have no reason to doubt. GeoSearch[33] is not listed at all. It runs a web-scale reverse image search and reads the pages it finds, so its score measures what the web already knows about a photograph as much as what can be inferred from it. Its authors evaluate with leakage in mind, and we do not dispute their numbers; they answer a different question. Photo-benchmark rows, ours included, can overlap with the training data the field shares; we flag it rather than claim immunity. GWS15k is by far the hardest because it samples Street View near random city centres worldwide, so 92 percent of its locations never appear in the training corpus the field shares. MP-16[27] is excluded because it is that shared training corpus rather than a test set: no standard held-out split, so no bar to clear.

Geolocation across benchmarks

On the closed-system vendors

The obvious question is how this compares to the closed-system tools an agency already evaluates. The short answer is that it cannot be compared, and we would rather say so plainly than work around it.

A closed-system vendor publishes a worked example in place of an evaluation. It shows one photograph and one result: a radius or an address and a handful of ranked candidates. There is no dataset, sample size, accuracy rate or standard split behind it. A worked example is a screenshot of one good day, and there is nothing in it a buyer can check.

The same pages carry a second thing, further down and in smaller type: that the output is an investigative lead rather than a confirmed location, and that the similarity score ranks candidates without being a probability that any one of them is right. Those caveats are correct, and we would write them ourselves. They are why confidence here limits what a result is allowed to claim instead of sitting beside it as decoration.

But a closed system cannot hold both ends. A page cannot advertise meter-level resolution at the top and disclaim at the bottom that its ranking carries no probability. One of those is the product and the other is the legal position, and an agency finds out which one it bought after it has signed. That is what closed means in practice. Private weights are ordinary; private evidence is the problem.

So there is no closed-system column in the tables above. A worked example cannot be placed beside 21,131 images without flattering somebody dishonestly, and we would rather not do that in either direction. What can be compared is how each side reports. Every number here sits on a public split, at a stated sample size, against named methods with citations, including the splits where Manta does not lead. That is a claim an agency can audit before it signs, and it is the only difference that does not require taking a vendor at their word.

The full results go to organizations evaluating Manta, including the splits where it does not lead. The report carries the OSV-5M full-test row and where it sits against the current bar, every classical split with the competing methods beside it, the failure modes we know about, and the evaluation protocol we recommend for testing on your own material.

Request the technical report · Request organization access

YFCC4k update

Geo-Estimation’s new YFCC4k run is ahead at all five thresholds. The best-other comparator is selected per threshold, so the margin is against a different method at different distances.

ThresholdMantaBest otherLead
@1 km35.0433.00 (Pinpoint)+2.04
@25 km47.3944.40 (Pinpoint)+2.99
@200 km58.5057.50 (Pinpoint)+1.00
@750 km73.1571.80 (Pinpoint)+1.35
@2500 km84.5684.50 (Pinpoint)+0.06

The OSV-5M table uses Oceanir’s latest reported board and highlights the highest numerical score at each radius. Run-to-run variability matters when interpreting small margins. The median error comes from the same predictions, measured by a standalone scorer; Chipoint-2 reports its median rounded to the whole kilometre. Lead is Manta minus the comparison score; positive is better. Pinpoint’s YFCC4k curve is taken from its paper, which publishes every threshold.

Method@1 km@25 km@200 km@750 km@2500 km
Im2GPS3k[26]
2,997 images. The split most papers lead with.
Geo-Estimation (Manta)bar22.5250.8266.5780.9191.16
REVERSE[24]22.548.359.373.584.8
Pinpoint[3]20.547.463.579.090.2
Geo-ADAPT[37]17.945.362.677.989.5
GeoRanker[25]18.7945.0561.4976.3189.29
HierLoc[39]11.343.858.474.185.1
GeoSURGE[29]17.242.558.174.687.6
GeoMetric[38]17.742.256.871.786.1
Geo-R[43]18.1041.5358.3175.3386.42
DualGeo[35]17.2541.4755.7671.7185.05
Chipoint v2[30]16.541.254.170.784.8
G3[25]16.6540.9455.5671.2484.68
GLOBE[44]9.8440.1856.1971.4582.38
Img2Loc[25]15.3439.8353.5969.7082.78
GeoToken[36]16.839.653.870.885.0
GaGA (MP16-Pro)[42]15.037.149.567.382.4
PIGEON[28]11.336.753.872.485.3
LocDiff-H[34]15.336.556.475.287.4
GRE Suite[31]11.3035.3051.7069.3085.70
GeoCLIP[4]14.1134.4750.6569.6783.82
Concept-Aware[46]13.234.049.868.283.5
GaGA[42]11.733.048.067.182.1
GeoLocSFT[45]8.8032.7047.2065.8082.25
StreetCLIP[48]—22.437.461.380.4
Pipelines over general-purpose models
A routing layer, tool loop or prompting stage wrapped around a general-purpose model such as Gemini or GPT-4V, rather than geolocation weights of its own. Shown in full and not ranked above, because each row measures the wrapper as much as the model. Their numbers are not in dispute.
GeoRouter[32]20.8250.4865.7380.3590.66
GeoToken + Gemini[36]19.046.060.176.688.8
VLM-VPR[41]18.6245.6559.7973.7185.95
SpotAgent[40]14.1240.3657.8073.4385.75
NAVIG[47]5.528.949.168.384.0
YFCC4k[26]
4,536 Flickr images. Noisier, more indoor and close-up.
Geo-Estimation (Manta)bar35.0447.3958.5073.1584.56
Pinpoint[3]33.044.457.571.884.5
GeoRanker[25]32.9443.5454.3269.7982.45
GeoMetric[38]31.241.151.867.881.0
Geo-ADAPT[37]32.539.155.470.884.5
DualGeo[35]27.4936.4545.0361.5875.92
G3[25]23.9935.8946.9864.2678.15
GeoToken[36]24.335.346.664.278.6
Chipoint v2[30]19.134.544.260.276.0
GaGA (MP16-Pro)[42]24.333.743.961.476.3
GeoSURGE[29]19.933.648.767.482.0
RFM S2 (10M)[49]—33.545.361.177.7
Img2Loc[25]19.7830.7141.4058.1174.07
HierLoc[39]8.430.243.361.775.8
REVERSE[24]14.127.538.153.870.6
PIGEON[28]10.423.740.662.277.7
RFM S2[49]—23.736.454.573.6
Geo-R[43]10.4722.6740.0460.8375.84
GaGA[42]6.918.934.556.771.6
GeoLocSFT[45]5.2118.5832.6453.1072.10
Single-threshold reports
Published at one distance only, so they cannot be ranked against the curves above.
GeoCLIP[4]10.0————
Pipelines over general-purpose models
A routing layer, tool loop or prompting stage wrapped around a general-purpose model such as Gemini or GPT-4V, rather than geolocation weights of its own. Shown in full and not ranked above, because each row measures the wrapper as much as the model. Their numbers are not in dispute.
GeoRouter[32]32.9846.0157.5272.0283.02
GeoToken + Gemini[36]25.438.551.468.081.0
SpotAgent[40]7.3021.5236.1855.0070.77
YFCC26k[26]
21,131 Flickr images. The larger YFCC split. Competitor curves are taken from GeoSURGE Table 4, which reproduces every prior method on this split under one protocol, except PIGEOTTO, which is shown at its own paper's headline result rather than the weaker ME16-only variant. Read the margin here with that in mind: the strongest recent retrieval and LVLM methods have not published on 26k. Pinpoint, GeoRanker, GeoRouter, REVERSE, G3, DualGeo and Chipoint v2 all report on YFCC4k, where the top of the field sits within about two points of us at 1 km, and none of them report 26k. LocDiff does, in its hybrid form, and is listed. So the field on this split is LocDiff, GeoSURGE and older work, and the double-digit gap is partly a statement about who has been measured rather than how far ahead we are. We would rather say that than let the two YFCC tables be lined up against each other and have it said for us.
Geo-Estimation (Manta)bar30.6845.3158.0573.4886.02
GeoMetric[38]20.333.746.464.179.3
GeoSURGE[29]17.831.545.164.379.3
RFM (S2-10M)[29]5.329.040.957.875.8
LocDiff-H[34]13.226.041.964.580.3
PIGEOTTO[28]10.525.842.763.279.0
GeoDecoder[28]10.123.934.149.669.0
GeoCLIP[4]11.622.236.757.576.0
TransLocator[28]7.217.828.041.360.6
GR/Qwen-VL[29]4.017.428.948.167.8
ISNs[28]5.312.319.031.950.7
PlaNet[29]4.411.016.928.547.7
GWS15k[28]
Street View shots near random city centres. The released split resolves to 14,955 images; Manta is scored against all of them, with the 394 that could not be scored counted as misses. 92 percent of its locations are unseen in the training corpus the field shares, which is why every number here collapses. Manta leads every method at all five thresholds, including LocDiff, the strongest published result below 2,500 km: 1 km (2.39 against LocDiff's 2.1), 25 km (15.23 against 12.4), 200 km (45.67 against 33.7), 750 km (78.00 against 67.0) and 2500 km (93.62 against PIGEOTTO's 85.1). The 1 km lead is 0.29 points, the narrowest of the five.
Geo-Estimation (Manta)bar2.3915.2345.6778.0093.62
LocDiff[34]2.112.433.767.085.0
PIGEOTTO[28]0.79.231.265.785.1
GeoSURGE[29]1.04.621.954.780.8
GRE Suite[31]0.94.118.954.878.3
GeoCLIP[4]0.63.116.945.774.1
GeoDecoder[28]0.71.58.726.950.5
TransLocator[28]0.51.18.025.548.3
ISNs[28]0.050.64.215.538.5
Pipelines over general-purpose models
A routing layer, tool loop or prompting stage wrapped around a general-purpose model such as Gemini or GPT-4V, rather than geolocation weights of its own. Shown in full and not ranked above, because each row measures the wrapper as much as the model. Their numbers are not in dispute.
VLM-VPR[41]0.4810.2832.6164.1485.25

Accuracy at each distance threshold, self-reported by each method unless noted. Groups are ordered by how many of the five thresholds each row tops, ties resolved at 25 km, the same measure the bar uses. “Bar” is derived, not assigned: within each split it marks whichever full-curve method tops the most of the five thresholds, with ties resolved at 25 km. Methods that publish a single threshold are listed separately and never badged, because a method cannot win a comparison it only partly enters. Rows are not perfectly like-for-like: methods differ in how they are built and trained. Nor are they equally checkable. Pinpoint posts the highest non-Manta figure on YFCC4k at 1 km (33.0). Its figures come first-hand from Virtualitics’ paper[3], which reports every threshold we quote, and its code is released for academic analysis. The trained weights are not, a choice the paper ties to misuse risk, so its numbers can be read at their source but not re-run by us or anyone else. We report them anyway, because leaving out the closest result to our own would be the more convenient omission, but a number you can read is not the same kind of evidence as a number you can re-run, and the row does not say so on its own. No method leads all four classical splits.

Street-level accuracy on YFCC26k, the largest of the classical photo splits at 21,131 images. Competitor rows are GeoSURGE Table 4, which reproduces every prior method under one protocol.Higher is better · accuracy at 1 km · Manta solid, every published comparator outlined · a method appears only if it reports this split.
YFCC4k at every threshold. The comparator is the strongest published result at that distance, so it changes method as the radius grows.Higher is better · each pair is Manta against the single strongest published number at that distance, so the outlined bar is not one method across the chart.
GWS15k against the three strongest methods that report every threshold: LocDiff, PIGEOTTO, which set the split, and GeoSURGE, which reruns it under its own protocol. Manta leads all three at every threshold, against PIGEOTTO’s headline result. All 14,955 images are scored; the 394 that could not be scored count as misses.Higher is better · both comparators are drawn at every threshold, so no result is dropped for being unflattering.

VPR benchmarks

Street-Match is the mode that resolves a photograph to the exact street or building. These are Manta’s results on seven standard VPR benchmarks, alongside every published model we could find numbers for. Recall@1 throughout.

What this table compares. Every row is a model: a specific set of trained weights evaluated end to end. It is not a comparison of techniques. Aggregation schemes, mining strategies and fine-tuning recipes are how these models are built, and several rows here share them, but a technique has no Recall@1 and cannot be put in a column. Where a paper reports one idea in several configurations we take its headline model rather than listing the idea once and the ablations beside it.

And where we sit in it. A ◇ marks a published result with no public code or weights. Street-Match carries one too. It is proprietary, we do not release its weights, and nobody outside Oceanir can reproduce our rows from scratch — so we are in exactly the category we are flagging, and it would be dishonest to mark competitors for it and exempt ourselves. What we can offer instead of weights is your own imagery: send the hardest cases you have and we will return what the system produces. Rows without a ◇ you can download and run today, and the citation links the weights. We keep the ◇ rows rather than dropping them, because two of them beat us and removing a competitor that wins is the more self-serving edit, not the more rigorous one.

A dash means the model was not evaluated on that benchmark in its paper or in subsequent reproductions.

ModelYearReportedSPEDPitts30kMSLS-valNordlandAmsterTimeTokyo 24/7SVOX
1SAGE (8448-D)[18]20266/798.995.894.596.083.597.5—
2Street-Match 1.0 ◇20267/795.3994.9594.6097.2174.2598.4398.74
3SAGE-L (no InteractHead)[18]20267/792.194.794.294.865.698.198.8
4BoQ[12]20247/792.593.793.890.663.098.199.0
5SuperPlace NVL-FT² (B) ◇[16]20257/787.593.794.391.462.396.898.6
6EMVP ◇[14]20246/794.694.093.988.765.696.8—
7SuperVLAD[13]20246/793.295.092.291.063.995.6—
8SALAD[11]20247/792.192.492.290.058.894.698.2
9FoL[17]20256/792.193.993.187.864.696.2—
10SALAD-CM[11]20246/789.592.694.295.657.896.8—
11SelaVPR †[9]20247/788.692.890.887.355.294.097.2
CricaVPR[10]20245/791.394.990.090.7—93.0—
MixVPR[7]20235/784.791.588.076.2—85.1—
EigenPlaces[8]20235/770.292.589.171.2—93.0—
CosPlace[6]20225/775.588.482.858.5—87.3—
NetVLAD[5]20165/770.281.953.16.4—60.6—
SuperPlace NVL-FT² (L) ◇[16]20253/7—94.194.5——97.1—
EffoVPR † ◇[15]20253/7—93.992.8——98.7—

All values are Recall@1. Our own rows are set in italics throughout this page: they are Oceanir runs, not figures reported by a third party. Street-Match 1.0 is evaluated one query image at a time on all seven benchmarks. † indicates a two-stage method with re-ranking. Published numbers sourced from the original papers and subsequent reproductions (SAGE, ICLR 2026; SuperPlace, 2025; BoQ, CVPR 2024). Where a paper reports several configurations we quote its headline one rather than mixing rows. SAGE is listed twice because the paper reports it two ways. Its InteractHead applies attention across a batch while building the training graph, and the paper states that this cost sits in training and leaves inference unaffected, so both rows are evaluated per image at inference, the same way Street-Match is evaluated. The second row is the paper’s Table 9 ablation, trained without that head. We show both rather than whichever suits us. Nordland uses the full summer-vs-winter protocol. MSLS-val uses 25 m with azimuth within 40 degrees. Manta’s Nordland evaluation differs slightly in frame alignment. MSLS-val is scored over the published 740 queries. SVOX protocol is unverified. AmsterTime is the cross-era archival-to-modern matching task that our product solves daily, and the benchmark we care about most.

Seven benchmarks in, seven reported, Street-Match 1.0 ranks second of the eleven models that report at least six. We are saying second because second is what the numbers say. The model above us is SAGE at 8448 dimensions[18]: peer reviewed at ICLR 2026, weights released under MIT, evaluated per image the same way we are, and it beats us on five of seven boards. A shipping product placing second to the current state of the art, on every benchmark the field uses, is a result we are content to publish unedited. First place on a table we drew ourselves would be worth less.

Nordland is the one board where we lead SAGE, and the only margin in this row wide enough to mean anything.97.21 against its 96.0, and against SALAD-CM’s 95.6[11], over 27,592 queries, where the 95 percent band is 0.21 points. A lead six times its own noise floor on the largest split in the table is a real result, and it is a lead over the model at the top of the table rather than over the field with that model set aside.

The rest of the row is read against the same test the OSV-5M table uses: a margin counts only when it clears the sampling band for that split’s query count. Three boards are inside it and we record them as ties. MSLS-val[20] is 94.60 against SAGE’s 94.5, but 740 queries put the band at 1.64 points, so a tenth of a point decides nothing. Tokyo 24/7[22] is 98.43 against EffoVPR’s 98.7[15] on 315 queries, and SVOX is 98.74 against BoQ’s 99.0[12] on 823. Small splits cannot separate the top of this table from itself.

Three are separable losses and we are not going to round them off. Pitts30k[23] is 94.95, 0.85 behind SAGE. SPED is 95.39, 3.51 behind. AmsterTime[19] is 74.25, clear of every other model in the table but 9.25 behind SAGE’s 83.5. AmsterTime is the cross-era archival-to-modern task our product does every day, so the board we care about most is the one we trail by the widest margin.

Against the full field that is one win, three ties and three losses. SAGE leads five of the seven boards and the table overall, 94.37 against our 93.37. More on the work in the Street-Match post.

Limits

Interiors give it very little. Overcast skies remove the light it reads latitude from. A stretch of motorway in one temperate country looks a great deal like a stretch of motorway in another, and there are photographs where the honest answer is a region and a shrug. We would rather say so here than have you find out in front of somebody who was relying on the number.

Manta answers where a photograph was taken. It is not built to find people and we do not permit it to be used that way. That is written into our terms rather than left to good intentions. The question worth answering is whether an image is what it claims to be, and answering it does not require pointing anything at anybody.

The Manta suite

Manta is a single model. It does two jobs, each a mode that answers a different question and fails differently. Geo-estimation says roughly where on Earth. Street-Match says this exact place. You choose the mode that fits the work.

Street-Match resolves a photograph to the exact street or building. The VPR benchmark results above are Manta’s, running in Street-Match mode. It is in research preview now.

How it ships

Manta is a different kind of system than anything we have shipped before, built to commit to a claim and hold ground on it rather than offer a plausible answer and move on. A model that behaves that way earns its rollout in stages.

Enterprise organizations first, then selected pilots and research partners, then wider access. Controlled access is a stage, not a destination. For now, access is granted to an organization, not a person: we verify the organization, it names the people who will use it and what they will use it for, and the grant is scoped to that. Coverage is enabled per city against a declared purpose rather than switched on globally, because “which cities can you do” and “what are you doing with them” are the same question. If a request cannot survive being written down, it does not get provisioned.

We are starting with organizations that can put Manta inside a reviewed workflow, a compliance record, an investigation, a case file with someone accountable for what it concludes. We do not yet know how it behaves across the full range of images people submit outside that setting, and we would rather learn on our own terms than find out from the outside.

If the capability gives you pause, you have read it correctly. It gives us pause too. That is why it ships this way, and it is no reason to keep quiet about it.

Availability

Manta is available now to select organizations under a reviewed enterprise workflow. Orca remains the default across the web app and the API for everyone else, and existing integrations continue to work unchanged. Individuals keep Orca, which answers where a photograph was taken and stops there.

What we learn from Manta is what moves into Orca, so the improvement reaches everyone without carrying over the parts that are still unproven.

Manta is not self-serve, and the web app runs Orca. If you want to see how Manta handles your material, send the hardest cases you have, including ones you already know the answer to. Request organization access · [email protected]

Go deeper: Street-Match · Newsroom case study

References

49references — tap ↵ to jump back to the passage that cites them.

  1. [1]

    Astruc et al.. OpenStreetView-5M: The Many Roads to Global Visual Geolocation ↗. CVPR 2024. ↵

    The open-world geolocation benchmark used for Manta's headline numbers: 5.1 million street-level images across 225 countries, with accuracy reported at 25 km, 200 km, 750 km and 2500 km thresholds.

  2. [2]

    Chiikabu Labs. Chipoint (v1) ↗. Hugging Face. ↵

    State of the art on OSV-5M until Chipoint v2 superseded it; Chipoint v2 is the current bar. Manta's current OSV-5M board (45.77 @25 km, 81.83 @200 km, 92.57 @750 km, 97.30 @2500 km) is above this baseline at every radius. Self-reported: 37.63 @25 km, 70.56 @200 km, 84.94 @750 km, 93.36 @2500 km, GeoScore 4174, mean error 716 km, median 51 km, country accuracy 85.55, admin-1 accuracy 61.7. CC-BY-NC-4.0. Its card states that Chipoint could not evaluate on im2gps, im2gps3k, YFCC4k or GWS15k because those splits are private or unretrievable.

  3. [3]

    Chuzhoy, Hu, Arora, Ro, Sahu. Pinpoint: Grounded Worldwide Image Geolocation via Cross-Source Retrieval and Reranking ↗. Virtualitics · arXiv 2606.04133, June 2026. ↵

    Virtualitics' retrieve-and-rerank model, trained on both Flickr photos and street-view imagery. Every figure we quote is first-hand from the paper: OSV-5M (Table 3) 35.6 @25 km, 67.5 @200 km, 83.7 @750 km, 93.2 @2500 km, GeoScore 4114, mean error 743 km; Im2GPS3k (Table 1) 20.5, 47.4, 63.5, 79.0, 90.2 and YFCC4k (Table 1) 33.0, 44.4, 57.5, 71.8, 84.5 at 1, 25, 200, 750 and 2500 km, with a YFCC4k median error of 84.9 km. The paper gives no OSV-5M median. Its code is released for academic analysis; the trained weights are not, so these results can be read at their source but have not been reproduced by Oceanir.

  4. [4]

    Vivanco Cepeda, Nayak, Shah. GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localization ↗. NeurIPS 2023. ↵

    Contrastive CLIP-style geolocation model whose OSV-5M numbers (21.5 @25 km, 52.1 @200 km) anchor the low end of the comparison.

  5. [5]

    Arandjelović et al.. NetVLAD: CNN Architecture for Weakly Supervised Place Recognition ↗. CVPR 2016. ↵

    The original learned VPR architecture; its Pitts30k and Tokyo 24/7 numbers are the baseline every subsequent model in the table is measured against. Weights: github.com/Nanne/pytorch-NetVlad.

  6. [6]

    Berton, Masone, Caputo. Rethinking Visual Geo-localization for Large-Scale Applications ↗. CVPR 2022. ↵

    Introduces CosPlace and the SF-XL training set; source of the CosPlace row across SPED, Pitts30k, MSLS-val, Nordland and Tokyo 24/7. Weights: github.com/gmberton/CosPlace (also torch.hub).

  7. [7]

    Ali-bey, Chaib-draa, Giguère. MixVPR: Feature Mixing for Visual Place Recognition ↗. WACV 2023. ↵

    MLP-mixer aggregation over CNN features; source of the MixVPR row. Weights: github.com/amaralibey/MixVPR.

  8. [8]

    Berton et al.. EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition ↗. ICCV 2023. ↵

    Class-orientation-aligned training; source of the EigenPlaces row. Weights: github.com/gmberton/EigenPlaces (also torch.hub).

  9. [9]

    Lu et al.. Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition ↗. ICLR 2024. ↵

    Foundation-model adaptation for VPR with a re-ranking stage, marked † in the table; source of the SelaVPR row including its SVOX result. Weights: github.com/Lu-Feng/SelaVPR.

  10. [10]

    Lu et al.. CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition ↗. CVPR 2024. ↵

    Cross-image correlation attention; source of the CricaVPR row. Weights: github.com/Lu-Feng/CricaVPR.

  11. [11]

    Izquierdo, Civera. Optimal Transport Aggregation for Visual Place Recognition ↗. CVPR 2024. ↵

    SALAD's optimal-transport aggregation. The SALAD-CM row is SALAD retrained with CliqueMining, a separate paper by the same authors (Izquierdo and Civera, "Close, But Not There: Boosting Geographic Distance Sensitivity in Visual Place Recognition", ECCV 2024, arxiv.org/abs/2407.02422, weights at github.com/serizba/cliquemining); it holds the Nordland lead (95.6) cited in the prose.

  12. [12]

    Ali-bey, Chaib-draa, Giguère. BoQ: A Place is Worth a Bag of Learnable Queries ↗. CVPR 2024. ↵

    BoQ's 98.1 on Tokyo 24/7 is from BoQ's own paper (its Table 3, R@1). Note a weaker BoQ configuration scoring 91.1 on Tokyo 24/7 circulates via EffoVPR's comparison table; the 98.1 here is BoQ's headline result. Weights: github.com/amaralibey/Bag-of-Queries (torch.hub).

  13. [13]

    Lu et al.. SuperVLAD: Compact and Robust Image Descriptors for Visual Place Recognition ↗. NeurIPS 2024. ↵

    NetVLAD without cluster centres, using very few clusters plus discarded ghost clusters; source of the SuperVLAD row. Weights at github.com/Lu-Feng/SuperVLAD. This is a DIFFERENT paper from SuperPlace [16] despite sharing a first author.

  14. [14]

    Qiu et al.. EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing ↗. NeurIPS 2024. ↵

    Centroid-Free Probing over a frozen foundation backbone with Dynamic Power Normalization; a PEFT pipeline tailored to VPR. Holds the SPED lead (94.6) cited in the prose breakdown. NOTE: EMVP released no code, and its AmsterTime and Tokyo 24/7 figures here are marked as reproductions in SAGE Table 3 rather than reported by EMVP itself. No public repository or weights exist for EMVP as of 2026-09-16.

  15. [15]

    Tzachor et al.. EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition ↗. ICLR 2025. ↵

    No public repository or weights; paper-reported only. Two-stage method (†). Its 98.7 is the top score on Tokyo 24/7; Street-Match 1.0's 98.43 sits inside that split's sampling band, so the prose records it as a tie.

  16. [16]

    Lu et al.. SuperPlace: The Renaissance of Classical Feature Aggregation for Visual Place Recognition in the Era of Foundation Models ↗. arXiv 2025. ↵

    NVL-FT2 is SuperPlace's NetVLAD-Linear aggregation trained with its secondary fine-tuning strategy: NetVLAD learns a high-dimensional descriptor which one linear layer then compresses. (B) and (L) are ViT-B and ViT-L backbones of the same model, not separate publications. The SuperPlace paper states no code availability and we could find no public repository or weights, so these rows are paper-reported only. The L configuration's 94.5 on MSLS-val equals SAGE's; Street-Match 1.0's 94.60 is inside that split's 1.64-point band, a tie.

  17. [17]

    Wang et al.. Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition ↗. AAAI 2025. ↵

    Two-stage retrieve-then-rerank model guided by reliable discriminative local regions; weights at github.com/chenshunpeng/FoL. Source of the FoL row.

  18. [18]

    Chen et al.. SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition ↗. ICLR 2026. ↵

    Frozen DINOv2 backbone with parameter-efficient fine-tuning, single-stage retrieval, MIT licensed with weights released. Two rows are shown. SAGE (8448-D) is the paper's headline ViT-L configuration from Tables 2 and 3; it holds the AmsterTime lead (83.5), the SPED lead (98.9) and the Nordland lead (96.0). The second is Table 9, the same ViT-L trained without the InteractHead. Both are evaluated at 322x322, which is SAGE's standard inference resolution throughout. The InteractHead applies attention across a batch while building the training graph, which the paper reconstructs each epoch; the paper states that this cost is confined to training and that inference is unaffected, so SAGE is a single-stage per-image encoder at deployment and is directly comparable to the other rows. SAGE is built on the EMVP framework, which its authors reproduced because EMVP released no code. Weights: github.com/chenshunpeng/SAGE.

  19. [19]

    Yildiz et al.. AmsterTime: A Visual Place Recognition Benchmark Dataset for Severe Domain Shift ↗. IJCV 2022. ↵

    Archival-to-modern street matching in Amsterdam; the benchmark the article singles out as the one the product solves daily and is least satisfied with.

  20. [20]

    Warburg et al.. Mapillary Street-Level Sequences: A Dataset for Lifelong Place Recognition ↗. CVPR 2020. ↵

    Source of the MSLS-val protocol (25 m threshold, azimuth within 40 degrees) described in the table footnote.

  21. [21]

    NRK / SINTEF. The Nordland Dataset. ↵

    Four-season train journey recordings; the table footnote describes the full summer-vs-winter protocol used.

  22. [22]

    Torii et al.. 24/7 Place Recognition by View Synthesis ↗. CVPR 2015. ↵

    Day-night street-level matching in Tokyo, 315 queries. Street-Match 1.0 scores 98.43, behind EffoVPR's 98.7 and ahead of BoQ and SAGE-L at 98.1; every gap is inside the sampling band, so the prose records a tie.

  23. [23]

    Torii et al.. Visual Place Recognition with Repetitive Structures ↗. CVPR 2013. ↵

    Pittsburgh street-view dataset. Street-Match 1.0 scores 94.95, 0.85 behind SAGE's 95.8: one of the three separable losses discussed in the prose.

  24. [24]

    Li, Jia, Yin, Rong, Rao, Lyu, Zhang. REVERSE: Reinforcing Evidence Verification and Search for Agentic Image Geo-localization ↗. arXiv 2026. ↵

    Current bar on Im2GPS3k at 25 km (48.3), shown in the other-geolocation-benchmarks table.

  25. [25]

    Jia et al.. GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization ↗. NeurIPS 2025. ↵

    Current verified bar on YFCC4k (43.54 @25 km, 32.94 @1 km, 54.32 @200 km); also reports 45.05 @25 km on Im2GPS3k.

  26. [26]

    Vo, Jacobs, Hays. Revisiting IM2GPS in the Deep Learning Era ↗. ICCV 2017. ↵

    Defines the Im2GPS3k (2,997 images) and YFCC4k (4,536) test splits drawn from Flickr and the Yahoo Flickr Creative Commons 100M corpus.

  27. [27]

    Larson, Soleymani, Gravier, Ionescu, Jones. The Benchmarking Initiative for Multimedia Evaluation: MediaEval ↗. IEEE MultiMedia 2017. ↵

    Source of MP-16, the 4.7M-image MediaEval Placing corpus used as training data throughout the field. It has no standard held-out test split, which is why no bar is quoted.

  28. [28]

    Haas, Skreta, Alberti, Finn. PIGEON: Predicting Image Geolocations ↗. CVPR 2024. ↵

    Introduces GWS15k, built by sampling countries in proportion to surface area and pulling Street View within 5 km of a random city centre, so 92 percent of its locations are unseen in MP-16 training data. PIGEOTTO's headline GWS15k result, the ME16 + Landmarks model in the paper's main results table, is 0.7 @1 km, 9.2 @25 km, 31.2 @200 km, 65.7 @750 km and 85.1 @2500 km (415.4 km median), and that is the row we compare against. The supplement (Table 8) also reports an ME16-only ablation at 0.1, 8.7, 30.1, 64.0 and 84.7. GeoSURGE daggers PIGEOTTO here as a different random sample and excludes it, so both are shown rather than one.

  29. [29]

    Daruna et al.. GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings ↗. CVPR 2026. ↵

    Reports 42.5 on Im2GPS3k and 33.6 on YFCC4k, below REVERSE and GeoRanker respectively.

  30. [30]

    Chiikabu Labs. Chipoint v2 ↗. Hugging Face. ↵

    Successor to Chipoint (v1, July 2026). Self-reported on the OSV-5M test set (210,122 images): 2.66 @1 km, 44.83 @25 km, 81.1 @200 km, 92.11 @750 km, 96.82 @2500 km, GeoScore 4485, mean error 383 km, median 32 km. MIT licensed, with a street-view arm and a general-photo arm. Its general-photo arm is separately reported on Im2GPS3k (16.5 @1 km, 41.2 @25 km, 54.1 @200 km, 70.7 @750 km, 84.8 @2500 km) and YFCC4k (19.1, 34.5, 44.2, 60.2, 76.0 at the same thresholds). At 1 km it is fourth on Im2GPS3k and sixth on YFCC4k. The street-view arm generalises poorly off its own domain: reranked onto im2gps3k it scores 11.68 @25 km against the photo arm's 41.2.

  31. [31]

    GRE Suite authors. GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains ↗. arXiv 2025. ↵

    Reports 11.30 @1 km and 35.30 @25 km on Im2GPS3k, and 4.1 @25 km on GWS15k. Previously mis-cited here to the PIGEON paper; it is its own source.

  32. [32]

    Jia et al.. GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization ↗. arXiv 2026. ↵

    Routes each query between a retrieval paradigm and a generation paradigm with an LVLM backbone, trained on a distance-aware preference objective. Supplies the full YFCC4k comparison curve and leads Im2GPS3k at 750 km and 2500 km.

  33. [33]

    Le-Duc, Nguyen-Son, Dao. GeoSearch: Augmenting Worldwide Geolocalization with Web-Scale Reverse Image Search and Image Matching ↗. arXiv 2026. ↵

    Adds web-scale reverse image search to a retrieval-augmented pipeline around a large multimodal model, feeding it coordinates and text taken from the web pages it finds, with image matching and confidence gating to filter the noise. Reports Im2GPS3k and YFCC4k under a leakage-aware evaluation, with code and data public. Not listed in the tables because a web search answers a different question than a set of weights does.

  34. [34]

    Wang, Liu, Zhang, Zhou, Cao, Wu, Mu, Song, Xie, Lao, Mai. LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert Space ↗. NeurIPS 2025. ↵

    Table 1 (GeoCLIP evaluation setup) reports its L=47 model on GWS15k at 2.1 @1 km, 12.4 @25 km, 33.7 @200 km, 67.0 @750 km and 85.0 @2500 km: the strongest published result at every threshold but 2,500 km, where PIGEOTTO's 85.1 edges it. The hybrid LocDiff-H (L=23, LocDiff proposals with GeoCLIP retrieval restricted to 200 km around them) scores lower on this split (0.9 @1 km) but is its best row on the photo splits: Table 1 gives 15.3 / 36.5 / 56.4 / 75.2 / 87.4 on Im2GPS3k and 13.2 / 26.0 / 41.9 / 64.5 / 80.3 on YFCC26k at 1, 25, 200, 750 and 2500 km. Its YFCC4k figures (Table 2) use a different protocol, trained on OSV-5M and tested zero-shot, so they are not listed. GWS15k results appear only in the NeurIPS version, not the first arXiv release.

  35. [35]

    DualGeo authors. DualGeo: A Dual-View Framework for Worldwide Image Geo-localization ↗. arXiv 2026. ↵

    Two stages: image and semantic-segmentation features fused by cross-attention and aligned with GPS by dual-view contrastive learning, then geographic-cluster re-ranking of the retrieved candidates and a final coordinate prediction by a large multimodal model (Qwen3-VL-Plus). Every figure we quote is first-hand from Table I: Im2GPS3k 17.25, 41.47, 55.76, 71.71, 85.05 and YFCC4k 27.49, 36.45, 45.03, 61.58, 75.92 at 1, 25, 200, 750 and 2500 km. Its code is public on GitHub without a stated licence; we have not reproduced the results. It reports Im2GPS and its own few-shot set as well, which are not listed here.

  36. [36]

    GeoToken authors. GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction ↗. ICDM 2025 · arXiv 2511.01082. ↵

    Predicts a hierarchy of S2 cell tokens from coarse to fine. Every figure is first-hand: Table I, without a language model, Im2GPS3k 16.8, 39.6, 53.8, 70.8, 85.0 and YFCC4k 24.3, 35.3, 46.6, 64.2, 78.6 at 1, 25, 200, 750 and 2500 km. Table II adds a Gemini 2.0 Flash stage (free-generation mode): Im2GPS3k 19.0, 46.0, 60.1, 76.6, 88.8 and YFCC4k 25.4, 38.5, 51.4, 68.0, 81.0. The tables list that second configuration as a pipeline because a general-purpose model makes the final call.

  37. [37]

    Geo-ADAPT authors. Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models ↗. arXiv 2603.13628, March 2026. ↵

    A reasoning model trained with reinforcement learning to spend more or less reasoning depending on how locatable an image is; it uses no retrieval database at inference. Table 1, Geo-ADAPT-8B: Im2GPS3k 17.9, 45.3, 62.6, 77.9, 89.5 and YFCC4k 32.5, 39.1, 55.4, 70.8, 84.5 at 1, 25, 200, 750 and 2500 km. Its own baseline rows are re-runs and differ slightly from the figures those methods published; the figures quoted here are only its own.

  38. [38]

    GeoMetric authors. GeoMetric: Injecting Metric Geographic Structure into Worldwide Image Geo-Localization ↗. arXiv 2606.08918, June 2026. ↵

    Retrieval with a location-attention GPS encoder, then a large multimodal model (Qwen-VL-Plus) on the retrieved candidates; an earlier version of the paper circulated as TransGeoCLIP. Every figure is first-hand, with the multimodal stage: Table 1 Im2GPS3k 17.7, 42.2, 56.8, 71.7, 86.1 and YFCC4k 31.2, 41.1, 51.8, 67.8, 81.0; Table 2 YFCC26k 20.3, 33.7, 46.4, 64.1, 79.3. Without that stage (Table 3) Im2GPS3k is 14.5, 35.1, 48.8, 68.9, 82.1.

  39. [39]

    HierLoc authors. HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation ↗. arXiv 2601.23064, January 2026. ↵

    Aligns images with a hierarchy of country, region, subregion and city entities in hyperbolic space. Table 2 (trained on MediaEval’16): Im2GPS3k 11.3, 43.8, 58.4, 74.1, 85.1 with a 73.4 km median, and YFCC4k 8.4, 30.2, 43.3, 61.7, 75.8 with a 341.9 km median. Its OSV-5M table reports administrative-level classification accuracy rather than distance thresholds, so it is not comparable to the OSV-5M table above and is not listed there.

  40. [40]

    SpotAgent authors. SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning ↗. arXiv 2602.09463, February 2026. ↵

    A fine-tuned vision-language model that can call web search, map, geocoding and image tools while it reasons. Table 1: Im2GPS3k 14.12, 40.36, 57.80, 73.43, 85.75 and YFCC4k 7.30, 21.52, 36.18, 55.00, 70.77 at 1, 25, 200, 750 and 2500 km. Listed as a pipeline because part of its answer comes from external tools, not from weights alone.

  41. [41]

    VLM-Guided VPR authors. VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization ↗. arXiv 2507.17455, July 2025. ↵

    A vision-language model (GPT-4V or Gemini 1.5 Pro) picks a geographic submap, a place-recognition model retrieves within it, and the model re-ranks. Table 2, its best configuration: Im2GPS3k 18.62, 45.65, 59.79, 73.71, 85.95; GWS15k 0.48, 10.28, 32.61, 64.14, 85.25. Listed as a pipeline because a general-purpose model makes the final call.

  42. [42]

    GaGA authors. Towards Interactive Global Geolocation Assistant ↗. arXiv 2412.08907. ↵

    A geolocation-tuned multimodal model with a clue-and-dialogue dataset. The standard-benchmark figures are in its appendix: Im2GPS3k 11.7, 33.0, 48.0, 67.1, 82.1 and YFCC4k 6.9, 18.9, 34.5, 56.7, 71.6; with an MP16-Pro retrieval tool, Im2GPS3k 15.0, 37.1, 49.5, 67.3, 82.4 and YFCC4k 24.3, 33.7, 43.9, 61.4, 76.3. Its OSV-5M results are administrative-level accuracy, so they are not listed there.

  43. [43]

    Geo-R authors. Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach ↗. arXiv 2601.00388, January 2026. ↵

    A retrieval-free reasoning model trained with a distance-based reward. Table 1: Im2GPS3k 18.10, 41.53, 58.31, 75.33, 86.42 and YFCC4k 10.47, 22.67, 40.04, 60.83, 75.84 at 1, 25, 200, 750 and 2500 km.

  44. [44]

    GLOBE authors. Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models ↗. NeurIPS 2025 · arXiv 2506.14674. ↵

    GLOBE-7B, a reasoning model trained with reinforcement learning on a 33K reasoning set. Table 2: Im2GPS3k 9.84, 40.18, 56.19, 71.45, 82.38 at 1, 25, 200, 750 and 2500 km. Its OSV-5M table uses a 3,000-image subset, so it is not listed on the OSV-5M table.

  45. [45]

    GeoLocSFT authors. GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models ↗. arXiv 2506.01277, June 2025. ↵

    A Gemma 3 27B model fine-tuned on about 2,700 samples, with no retrieval database. Tables: Im2GPS3k 8.80, 32.70, 47.20, 65.80, 82.25 and YFCC4k 5.21, 18.58, 32.64, 53.10, 72.10 at 1, 25, 200, 750 and 2500 km.

  46. [46]

    Concept-Aware authors. Towards Interpretable Geo-localization: a Concept-Aware Global Image-GPS Alignment Framework ↗. arXiv 2509.01910, September 2025. ↵

    Adds a geography-driven concept subspace to image and location embeddings. Table 1: Im2GPS3k 13.2, 34.0, 49.8, 68.2, 83.5 at 1, 25, 200, 750 and 2500 km, reported against its own GeoCLIP reproduction (10.8, 31.1, 48.7, 67.6, 83.2).

  47. [47]

    NAVIG authors. Navig: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization ↗. arXiv 2502.14638, February 2025. ↵

    A vision-language model with a reasoning stage and a map-search stage. Its main results are on GWS5k; Im2GPS3k is in its appendix (Table 13): 5.5, 28.9, 49.1, 68.3, 84.0 at 1, 25, 200, 750 and 2500 km. Listed as a pipeline because it relies on map search.

  48. [48]

    Haas, Alberti, Skreta. Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization ↗. arXiv 2302.00275. ↵

    StreetCLIP, zero-shot. Table 1, Im2GPS3k: 22.4 @25 km, 37.4 @200 km, 61.3 @750 km, 80.4 @2500 km. It does not report 1 km.

  49. [49]

    Dufour et al.. Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation ↗. arXiv 2412.06781. ↵

    Riemannian flow-matching geolocation. Its YFCC4k table reports RFM S2 at 23.7, 36.4, 54.5, 73.6 and the 10M-image model at 33.5, 45.3, 61.1, 77.7 at 25, 200, 750 and 2500 km, with no 1 km. Its OSV-5M table reports country, region and city accuracy, not distance thresholds, so it is not listed on the OSV-5M table.

In this article

  1. 01What changed
  2. 02Reading the picture
  3. 03A record, not just a result
  4. 04Benchmarks
  5. 05Limits
  6. 06The Manta suite
  7. 07How it ships
  8. 08Availability

Physical inference for the world.

Book a demoOpen the app

Oceanir

Oceanir works from what is visible in an image, not its metadata.

Product

  • MantaNew
  • Orca

Use Oceanir

  • Open the app
  • Photo location finder
  • Find location from a photo
  • Platform
  • Pricing

Resources

  • How it works
  • Documentation
  • API reference
  • Insights
  • Status

Company

  • About Oceanir
  • Contact
  • Security
  • Product boundaries
  • X (Twitter)

© 2026 Oceanir, LLC. All rights reserved.

PrivacyTermsCookiesDPASaaS agreement