NewsVisual positioning

AI-powered visual positioning can find you from just one photo.

MultiSet matched an ordinary photo against a 3D scan of Austin. The visible scene led it back to the camera’s exact position.

ShareXFacebookLinkedIn
A 3D scan of Austin with colored lines connecting photo features to their matched positions in the city
Photo features are matched to the same points inside a 3D scan of Austin.Image: Bilawal Sidhu and MultiSet via X

Bilawal Sidhu took several photos around Austin and tested them with MultiSet’s private visual positioning system (opens in a new tab). One came from a high-rise window. The system did more than recognize the skyline. It found the building, the height, and the exact balcony where the camera had been standing.

The buildings, roads, and skyline became clues that led back to the camera. The view itself became location data.

How it found the balcony

To understand how, it helps to think of visual positioning as GPS for cameras. First, a place is scanned and turned into a 3D map. A new photo can then be compared with that map to find where the camera is and which way it is pointing.

Sidhu’s demo makes that process visible. Colored lines connect details in each Austin photo to the same points inside the city scan. Once enough details match, the photos and video frames snap back into the places where they were captured.

What made it possible

That precision did not come from the photo alone. The test worked because that part of Austin had already been scanned and Sidhu could access the matching map. This is not a public tool that can identify every home from any photo.

That boundary matters. MultiSet says its maps are private to each customer account (opens in a new tab), so other developers cannot search them as one shared public world map.

For now, that limits how widely a photo can be matched. Even so, detailed maps are becoming easier to create. MultiSet can use phone LiDAR, 360-degree video, point clouds, Matterport scans, textured models, and Gaussian splats. As more places are mapped, more ordinary scenes could be recognized this way.

Why the view matters

The experiment changes how we should think about photo privacy. Removing GPS metadata may not be enough when the buildings, roads, windows, and skyline form their own visual fingerprint.

A skyline photo can seem anonymous on its own. When it is compared with a matching map, the same view could point to a home, hotel, or office, then narrow the location to a building, floor, and window.

The same matching process also has useful applications. It can keep digital directions attached to the right door or help a robot locate itself indoors. That is the tension at the heart of visual positioning. The match can make spatial apps and robots more capable, while the location it reveals remains sensitive.

Today, not every skyline photo exposes an address. Once a place has been mapped, the pixels in the view may be enough.

More from Capture & Create.

See all stories
An InSpatio demo showing source views beside a new view of a man reconstructed from one video

Spatial AI · 3 min read

InSpatio turns one video into a moving 4D Gaussian scene.

The research preview creates new camera angles from one video clip, without a multi-camera setup.

Justin Ryan holding smart glasses outside Meta Lab

Apple smart glasses · 2 min read

Apple may unveil its first smart glasses at WWDC 2027.

If Bloomberg’s report is right, developers could see them next June and customers could get them by the end of 2027.

Adobe Premiere timeline and video editor arranged as floating panels in visionOS

Apple Vision Pro apps · 3 min read

Adobe Premiere now edits spatial video on Apple Vision Pro.

Premiere’s new visionOS app can combine flat and spatial clips, place titles in depth, and export a spatial video without leaving Vision Pro.